CRDT: Overcoming Information Gain Bias in Healthcare Data Mining

CRDT: Correlation Ratio Based Decision Tree Model for Healthcare Data Mining

2016-10-01
Smita Roy, Samrat Mondal, Asif Ekbal, Maunendra Sankar Desarkar
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces CRDT (Correlation Ratio Based Decision Tree), a novel classification model tailored for healthcare data mining. It replaces the traditional Information Gain (IG) metric with Correlation Ratio (CR) to achieve SOTA-comparable performance while eliminating the inherent bias toward attributes with numerous distinct values.

TL;DR

The Correlation Ratio Based Decision Tree (CRDT) addresses a fundamental flaw in traditional ID3/C4.5 algorithms: the bias toward high-cardinality attributes. By utilizing a statistical Correlation Ratio (CR) for splitting, this model provides more robust classifications for healthcare datasets like Hepatitis and Liver disease, where feature significance is often overshadowed by the number of distinct values.

The "Cardinality Trap" in Medical Data

Information Gain (IG) is the bedrock of many Decision Tree models. However, it possesses an "Achilles' heel"—it mathematically prefers attributes with more distinct values. In a clinical setting, an attribute like "Patient ID" would yield the highest Information Gain because it creates perfectly pure (but useless) partitions.

Authors Roy et al. argue that healthcare data is uniquely susceptible to this. Medical records often mix binary flags (smoker/non-smoker) with high-variance categorical data (blood types, symptoms). IG-based trees frequently pick the "noisier" high-cardinality features, leading to models that fail to generalize on new patients.

Methodology: The Power of Correlation Ratio

The core innovation lies in the transition from Entropy to Correlation Ratio (CR).

1. Intuition Behind CR

Unlike IG, which looks at the reduction in uncertainty, CR looks at the dispersion of values. A significant attribute is one where the average value within a specific outcome class (e.g., "Hepatitis Positive") is remarkably different from the overall population average.

2. The Logic

The algorithm follows a classic recursive partitioning structure but calculates the CR for each attribute at every node.

  • Numerator: Dispersion among individual classes.
  • Denominator: Dispersion across the whole population.

The model chooses the attribute that maximizes this ratio, ensuring that the split is based on true class-relevance rather than just the "splitting power" of many distinct values.

CRDT Algorithm Overview (Note: Refer to Algorithm 1 and 2 in the paper for the recursive construction logic and the adaptation of CR for nominal attributes.)

Experimental Results: Where CRDT Shines

The researchers tested CRDT against IG across eight benchmark datasets from the UCI repository.

Key Findings:

  • Superiority in Small/Complex Datasets: CRDT outperformed IG in the Hepatitis dataset (73.78% vs 71.19%) and the Indian Liver Patient Dataset (ILPD). These datasets are characterized by features with differing numbers of distinct values.
  • Consistency: In datasets where attributes had uniform cardinalities (like Spect-heart), CRDT matched IG's performance (74.33%) exactly, proving it is a safe "drop-in" replacement.
  • Robustness: The results indicate that as the dataset becomes "messier" with varying attribute types, CRDT’s lack of bias becomes a significant advantage.

Performance Comparison Table (Note: Table V in the paper details the cross-validation results across all 8 healthcare benchmarks.)

Critical Insight: When to Use CRDT?

The paper suggests a "complementary" relationship. CRDT is not a "silver bullet" to replace all Decision Trees, but it is a superior choice when:

  1. The dataset contains a mix of categorical and discretized numerical data.
  2. Prior IG models show signs of overfitting on high-cardinality features.
  3. The sample size is relatively small (like the Hepatitis or Statlog datasets), where every split decision critically impacts the final accuracy.

Conclusion and Future Directions

The CRDT model provides a mathematically grounded solution to the bias inherent in Information Gain. By shifting the focus to Correlation Ratios, the authors have created a tool that respects the biological significance of medical features over their statistical distribution. Future work could involve scaling this to ensemble methods like Random Forests to see if the "CR-Forest" outperforms the standard IG-based versions.


Author Affiliations: Smita Roy (Central University of Bihar), Samrat Mondal & Asif Ekbal (IIT Patna).

Find Similar Papers

Try Our Examples

  • Search for recent papers that propose alternative splitting criteria for Decision Trees to solve the multi-valued attribute bias problem beyond Gain Ratio and Correlation Ratio.
  • Which seminal work first adapted the statistical Correlation Ratio for categorical data mining, and how does the CRDT algorithm's mathematical implementation differ from that origin?
  • Explore instances where Correlation Ratio-based feature selection has been applied to Deep Learning architectures or Ensemble methods like Random Forests for medical diagnosis.
Contents
CRDT: Overcoming Information Gain Bias in Healthcare Data Mining
1. TL;DR
2. The "Cardinality Trap" in Medical Data
3. Methodology: The Power of Correlation Ratio
3.1. 1. Intuition Behind CR
3.2. 2. The Logic
4. Experimental Results: Where CRDT Shines
4.1. Key Findings:
5. Critical Insight: When to Use CRDT?
6. Conclusion and Future Directions