Ensemble Learning: Navigating Credit Risk in Vietnam’s Banking Sector
Machine Learning-Based Empirical Investigation for Credit Scoring in Vietnam’s Banking
This paper presents a comparative investigation of ensemble learning models—LightGBM, CatBoost, and Random Forest—for credit scoring in the Vietnamese banking sector. Utilizing the Kalapa Credit Score dataset, the authors demonstrate that ensemble methods significantly outperform traditional single-model baselines, with Random Forest achieving a top-tier F1-Score of 0.83.
Executive Summary
TL;DR: This study provides an empirical benchmark for credit scoring in Vietnam, comparing modern ensemble methods against traditional statistical models. By optimizing Random Forest and Gradient Boosting architectures, the researchers achieved an F1-Score of 0.83, demonstrating that ensemble intelligence can effectively mitigate the risks of "bad debt" in emerging financial markets.
Context: This work represents one of the first systematic applications of advanced Machine Learning to the Vietnamese banking context, specifically addressing the challenges of the COVID-19 pandemic's impact on debt recovery.
The Challenge: Noisy, Imbalanced, and Sparse Data
In the wake of the pandemic, Vietnamese banks faced an urgent need to refine their credit assessments. However, the data reality is harsh:
- Extreme Imbalance: "Bad" labels account for only ~30% of the data.
- Information Sparsity: Over 100 features had missing data rates exceeding 50%.
- Security Constraints: Encrypted fields (to protect customer privacy) create a "semantic gap" that makes feature engineering difficult.
The authors argue that traditional models like SVM are too riddled with "inductive bias" to handle such noisy environments effectively.
Methodology: The Power of Ensembles
The core of the solution lies in a multi-stage preprocessing and model selection pipeline.
1. Feature Engineering and Selection
To handle the 195 original features, the authors implemented Target Permutation. By shuffling targets to create a "null importance" distribution, they could filter out noise and retain only the 33 most predictive attributes (Information Value > 0.2).
2. Architecture Comparison
The paper benchmarks three heavyweights of the tabular data world:
- LightGBM: Chosen for its efficiency and low memory footprint.
- CatBoost: Leveraged for its native handling of categorical variables without explicit encoding.
- Random Forest: Utilized as a stability benchmark that aggregates multiple decision trees to reduce variance.

Experimental Insights: AUC vs. Stability
While data scientists often chase the highest AUC (Area Under Curve), this paper provides a critical "Academic Reality Check."
The CatBoost vs. Random Forest Paradox:
- CatBoost achieved the highest AUC (0.82).
- Random Forest achieved the highest F1-Score (0.83).
The authors' deeper analysis reveals that while CatBoost might predict higher potential revenue at certain thresholds, its profit volatility is much higher. For a bank, a model that is slightly less "accurate" on average but more stable across different thresholds (Random Forest) is often safer for deployment.

Results & Discussion
The results confirm that ensemble models are vastly superior to single-algorithm approaches in this domain. Random Forest's ability to maintain a high Recall (0.81) and F1-Score (0.83) makes it the most "suitable and flexible solution" for the Kalapa dataset.
Takeaway and Future Outlook
This work sets a baseline for Vietnamese Fintech. The shift from manual heuristics to ensemble-based scoring can save banks significant capital by accurately identifying high-risk borrowers.
Future Work: The authors aim to explore Deep Learning and more sophisticated Feature Extraction techniques to further reduce training time and increase the granularity of credit ratings.
Critical Note: While the "Black Box" nature of Random Forest was noted as a drawback, in the context of Vietnamese banking, the trade-off for higher predictive stability appears to be a winning strategy.
