Boosting Efficiency: Handling Imbalanced Churn Prediction with RUSBoost and Feature Selection
Handling Imbalanced Data in Churn Prediction Using RUSBoost and Feature Selection (Case Study: PT.Telekomunikasi Indonesia Regional 7)
This paper presents a hybrid approach for telecommunications churn prediction by combining Random Under-Sampling Boost (RUSBoost) with Information Gain (IG) feature selection. Using real-world broadband data from PT. Telekomunikasi Indonesia, the method achieves an optimized churn model specifically designed for highly imbalanced datasets.
TL;DR
Predicting which customers will leave (churn) is a classic "needle in a haystack" problem. In this study, researchers from Telkom University tackle a dataset where only ~1.5% of customers churned. By combining Information Gain (IG) for feature selection and RUSBoost for ensemble learning, they managed to increase prediction accuracy (F-Score) by 16% while slashing processing time by nearly half (48%).
The Challenge: The Imbalance Paradox
In the telecommunications industry, retaining an existing customer is significantly more profitable than acquiring a new one. However, the data reflects a harsh reality: churners are a tiny minority.
Standard machine learning models are designed to maximize overall accuracy. If 99% of your customers stay, a model can be 99% "accurate" by simply predicting that no one will ever leave. This makes the model useless for the business. This paper addresses two critical bottlenecks:
- Class Imbalance: The overwhelming dominance of the "non-churn" class.
- High Dimensionality: Too many irrelevant features (billing details, usage patterns, service faults) that slow down training.
Methodology: A Hybrid Defense
The authors propose a structured pipeline that moves from raw data to a balanced, efficient classifier.
1. Temporal Data Selection
Before training, the authors asked: How many months of history do we actually need? They tested windows of 3, 6, 9, and 12 months. Surprisingly, 9 months of history provided the richest signals for the model, proving that more data isn't always better—it's about the right window of behavior.
2. Feature Selection via Information Gain
Using Information Gain, the authors filtered out features with low entropy. For instance, in the 9-month dataset, they reduced the feature count to just 7 key attributes (primarily focused on usage patterns). This dimensionality reduction is what eventually led to the massive 48% speedup.
3. The Power of RUSBoost
RUSBoost (Random Under-Sampling Boost) is the core engine here. Unlike over-sampling (which can lead to overfitting by duplicating data), RUSBoost randomly removes examples from the majority class during each iteration of the boosting process. By integrating this with the AdaBoost.M2 algorithm and C4.5 decision trees, the model focuses its "attention" on the hard-to-predict churners.

Experimental Insights & Results
The experiment tested various "target class distributions"—essentially adjusting how aggressive the under-sampling should be.
Key Findings:
- The Sweet Spot: The highest F-Score was achieved when the training data was rebalanced to a 10% churn rate. Pushing for a 50/50 split actually degraded the F-Score, likely because too much majority-class information was lost.
- The Efficiency Gain: By using Information Gain (IG), the model wasn't just faster; it was smarter. The IG-enhanced version consistently outperformed the base RUSBoost model.

| Metric | Improvement with IG |
|---|---|
| F-Score | +16.0% (Avg) |
| Processing Time | -48.0% (Avg) |
Critical Analysis & Conclusion
This work demonstrates that for industrial-scale churn prediction, "less is more." By identifying that Usage and Billing are the primary drivers of churn, and by using RUSBoost to force the model to learn from the minority class, PT. Telkom Indonesia can more effectively target at-risk subscribers.
Takeaway for Practitioners: If you are dealing with a minority class (fraud detection, medical diagnosis, churn), don't just throw data at a model.
- Reduce features first to eliminate noise and save time.
- Sample intelligently—RUSBoost remains a robust, computationally efficient alternative to more complex over-sampling methods like SMOTE.
Future Directions: While C4.5 was used as the weak learner here, the same framework could be tested with more modern gradient-boosted trees (XGBoost/LightGBM) or even neural networks to see if the F-Score can be pushed beyond the current benchmarks.
