From Rankings to Recommendations: A Hybrid DEA-Naïve Bayes Approach to Regional Competitiveness
Predicting Regional Development Competitiveness Index Using Naive Bayes: Basis for Recommender System
This paper presents a regional development prediction model specifically for Region 1 of the Philippines, combining Data Envelopment Analysis (DEA) and the Naïve Bayes algorithm. By analyzing the Cities and Municipalities Competitiveness Index (CMCI), the system provides a 90.64% accuracy rate in predicting Local Government Unit (LGU) competitiveness across four pillars: Economic Dynamism, Government Efficiency, Infrastructure, and Resiliency.
Executive Summary
TL;DR: This study moves beyond mere performance tables by introducing a predictive recommender system for Local Government Units (LGUs) in the Philippines. By merging Data Envelopment Analysis (DEA) for efficiency benchmarking with Naïve Bayes for probabilistic prediction, the authors provide a framework that tells LGUs not just where they stand, but how much they need to improve to become competitive.
Positioning: This work serves as a practical bridge between data mining (CRISP-DM) and public policy making, specifically targeting the improvement of regional rankings in the ASEAN economic landscape.
Problem & Motivation: The Gap in Performance Metrics
Despite showing growth capacity, the Philippines' regional competitiveness has historically struggled with issues like infrastructure inadequacy and government inefficiency. The Cities and Municipalities Competitiveness Index (CMCI) provides a wealth of data across four pillars:
- Economic Dynamism
- Government Efficiency
- Infrastructure
- Resiliency
However, the raw index often fails to answer a critical executive question: "If we are ranked 10th, exactly how much more investment or policy reform is required to match the efficiency of the 1st ranked city?" This research addresses this by identifying the "Slack"—the quantifiable deficiency preventing a municipality from reaching 100% relative efficiency.
Methodology: The Analytical Twin Engine
The study follows the CRISP-DM (Cross-Industry Standard Process for Data Mining) lifecycle. The technical core lies in the two-step mathematical evaluation:
1. Data Envelopment Analysis (DEA)
DEA is used to determine Relative Efficiency. LGUs are treated as Decision Making Units (DMUs). If Dagupan City scales the highest, it is set as the 100% benchmark. Every other LGU's score is divided by Dagupan's score to find its efficiency gap.
2. Naïve Bayes Classification
Once the efficiency is ranked, the top 10 LGUs are labeled as "Competitive" (Yes/No). The Naïve Bayes algorithm then processes the conditional probabilities of being competitive based on the LGU category and indicator performance.
Figure 1: The integration of CMCI framework with DEA and Naïve Bayes for decision support.
Experiments and Results
The study analyzed 125 LGUs in Region 1. Using the WEKA tool, the Naïve Bayes algorithm demonstrated an impressive 90.64% accuracy in predicting competitiveness.
The "Vigan Case Study"
In the Economic Dynamism pillar, Vigan City was found to have a relative efficiency of 60.82% compared to Dagupan. The model calculated a Slack of 0.0067, meaning Vigan must increase its indicator score by 64.42% to reach the "Efficiency Frontier" established by Dagupan.
Table 1: DEA calculation showing Relative Efficiency, Targets, and Slack Percentages for Region 1 cities.
Probabilistic Insights
The Naïve Bayes classifier further revealed that for 1st-2nd class municipalities:
- If Local Economy Size and Growth are both "Yes" (high), there is an 82.15% likelihood of being overall competitive.
- This allows the "Recommender System" to forecast that LGUs falling into these categories should expect a competitive status, provided they maintain current growth trajectories.
Critical Insight & Conclusion
Takeaway
The true value of this paper isn't just the 90.64% accuracy; it is the prescriptive nature of the results. By providing a "Target" score, the model transforms abstract rankings into actionable KPIs for local mayors and regional planners.
Limitations & Future Work
While the model is robust, it assumes Class Conditional Independence—meaning it treats Government Efficiency and Infrastructure as independent, whereas in reality, they are often deeply coupled. Future research could explore Bayesian Networks (which account for dependencies) or Random Forests to see if the predictive accuracy can be pushed beyond 95%.
Ultimately, this study proves that data mining is not just for tech companies; it's a vital tool for building more resilient and competitive government institutions.
