Precision Marketing: Solving Class Imbalance with Tomek Link and LightGBM
Data Mining Techniques in Direct Marketing on Imbalanced Data using Tomek Link Combined with Random Under-sampling
This paper introduces a hybrid data mining framework for direct marketing, combining Tomek Link and Random Under-sampling (T-link/RUS) to address severe class imbalance. By integrating Chi-squared feature selection and a tuned LightGBM classifier, the authors achieved an industry-leading 0.947 sensitivity and 0.896 ROC AUC on the UCI bank marketing dataset.
TL;DR
Direct marketing is a "needle in a haystack" problem. This paper presents a refined pipeline that combines Tomek Link (noise removal) with Random Under-sampling and LightGBM to identify potential customers. The result? A massive jump in sensitivity (from 56% to 94%) and a significant reduction in computational overhead.
Background: The Cost of Missing a "Yes"
In bank telemarketing, the cost of misclassifying a potential subscriber (False Negative) is much higher than calling someone who isn't interested (False Positive). However, standard machine learning models are "lazy"—they achieve high accuracy by simply predicting "No" for everyone. To solve this, we must address the Class Imbalance and the Overlapping Boundary problem.
Methodology: Cleaning the Decision Boundary
The authors' core insight is that not all majority-class samples are created equal. Some "No" samples are located right in the middle of "Yes" clusters, acting as noise.
1. Hybrid Resampling (T-Link + RUS)
Instead of just deleting data at random, the paper uses Tomek Link. A Tomek link exists if two samples of different classes are each other's nearest neighbors. Removing the majority-class sample from such a pair "clears" the boundary, making it easier for models like XGBoost or LightGBM to find the optimal hyperplane.
2. Feature Refinement
Using the Chi-squared method, the authors reduced the feature count from 63 to 22. This wasn't just about speed; it removed non-informative noise (like 'housing' or 'loan' status in this specific dataset) that was degrading predictive power.

Experimental Results: SOTA Performance
The combination of T-link/RUS and Chi-squared selection was tested across 10 different classifiers.
- LightGBM Winner: LightGBM emerged as the champion, benefiting from its leaf-wise growth strategy which works exceptionally well on cleaned, balanced data.
- Metrics: Sensitivity reached 0.947, meaning the model catches nearly 95% of all potential customers.
- Efficiency: Training times plummeted. For Gradient Boosting, training time dropped from 477 seconds to a mere 19 seconds.

Critical Insight: Why This Works
Most practitioners go straight to SMOTE (Oversampling). However, this paper shows that for marketing data, Under-sampling (RUS) combined with Noise Removal (Tomek) often outperforms SMOTE. This is because SMOTE can inadvertently create "synthetic" noise if the original minority samples are already outliers. By cleaning the majority class first, the model gains a "clearer view" of what a successful customer actually looks like.
Summary & Outlook
This research provides a production-ready blueprint for imbalanced tabular data:
- Clear the noise with Tomek Links.
- Balance the ratio with Random Under-sampling.
- Optimize the features with Chi-squared to save costs.
- Execute with a tuned LightGBM.
Future work could involve testing this pipeline on even larger datasets or exploring the use of Cost-Sensitive Learning in conjunction with these resampling techniques to further penalize False Negatives.
