Precision Marketing: Solving Class Imbalance with Tomek Link and LightGBM

Data Mining Techniques in Direct Marketing on Imbalanced Data using Tomek Link Combined with Random Under-sampling

2021-05-27
Ümit Yilmaz, Cengiz Gezer, Zafer Aydin, Vehbi Çagri Güngör
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a hybrid data mining framework for direct marketing, combining Tomek Link and Random Under-sampling (T-link/RUS) to address severe class imbalance. By integrating Chi-squared feature selection and a tuned LightGBM classifier, the authors achieved an industry-leading 0.947 sensitivity and 0.896 ROC AUC on the UCI bank marketing dataset.

TL;DR

Direct marketing is a "needle in a haystack" problem. This paper presents a refined pipeline that combines Tomek Link (noise removal) with Random Under-sampling and LightGBM to identify potential customers. The result? A massive jump in sensitivity (from 56% to 94%) and a significant reduction in computational overhead.

Background: The Cost of Missing a "Yes"

In bank telemarketing, the cost of misclassifying a potential subscriber (False Negative) is much higher than calling someone who isn't interested (False Positive). However, standard machine learning models are "lazy"—they achieve high accuracy by simply predicting "No" for everyone. To solve this, we must address the Class Imbalance and the Overlapping Boundary problem.

Methodology: Cleaning the Decision Boundary

The authors' core insight is that not all majority-class samples are created equal. Some "No" samples are located right in the middle of "Yes" clusters, acting as noise.

1. Hybrid Resampling (T-Link + RUS)

Instead of just deleting data at random, the paper uses Tomek Link. A Tomek link exists if two samples of different classes are each other's nearest neighbors. Removing the majority-class sample from such a pair "clears" the boundary, making it easier for models like XGBoost or LightGBM to find the optimal hyperplane.

2. Feature Refinement

Using the Chi-squared method, the authors reduced the feature count from 63 to 22. This wasn't just about speed; it removed non-informative noise (like 'housing' or 'loan' status in this specific dataset) that was degrading predictive power.

Proposed Approach Diagram

Experimental Results: SOTA Performance

The combination of T-link/RUS and Chi-squared selection was tested across 10 different classifiers.

  • LightGBM Winner: LightGBM emerged as the champion, benefiting from its leaf-wise growth strategy which works exceptionally well on cleaned, balanced data.
  • Metrics: Sensitivity reached 0.947, meaning the model catches nearly 95% of all potential customers.
  • Efficiency: Training times plummeted. For Gradient Boosting, training time dropped from 477 seconds to a mere 19 seconds.

Performance Comparison Table

Critical Insight: Why This Works

Most practitioners go straight to SMOTE (Oversampling). However, this paper shows that for marketing data, Under-sampling (RUS) combined with Noise Removal (Tomek) often outperforms SMOTE. This is because SMOTE can inadvertently create "synthetic" noise if the original minority samples are already outliers. By cleaning the majority class first, the model gains a "clearer view" of what a successful customer actually looks like.

Summary & Outlook

This research provides a production-ready blueprint for imbalanced tabular data:

  1. Clear the noise with Tomek Links.
  2. Balance the ratio with Random Under-sampling.
  3. Optimize the features with Chi-squared to save costs.
  4. Execute with a tuned LightGBM.

Future work could involve testing this pipeline on even larger datasets or exploring the use of Cost-Sensitive Learning in conjunction with these resampling techniques to further penalize False Negatives.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize advanced hybrid resampling techniques like SMOTE-Tomek or ADASYN-Tomek in financial customer churn or marketing prediction.
  • Who first proposed the Tomek Link method, and how has its implementation evolved for high-dimensional categorical data since its inception in 1976?
  • Explore studies that apply LightGBM and Chi-squared feature selection to other imbalanced tabular datasets such as fraud detection or medical diagnosis.
Contents
Precision Marketing: Solving Class Imbalance with Tomek Link and LightGBM
1. TL;DR
2. Background: The Cost of Missing a "Yes"
3. Methodology: Cleaning the Decision Boundary
3.1. 1. Hybrid Resampling (T-Link + RUS)
3.2. 2. Feature Refinement
4. Experimental Results: SOTA Performance
5. Critical Insight: Why This Works
6. Summary & Outlook