LIME, SHAP, or Anchors? Decoding the "Black Box" of Healthcare AI

Interpretability in healthcare: A comparative study of local machine learning interpretability techniques

2020-11-24
Radwa El Shawi, Youssef Sherif, Mouaz H. Al-Mallah, Sherif Sakr
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a comparative experimental evaluation of three prominent local model-agnostic interpretability techniques—LIME, SHAP, and Anchors—specifically within the healthcare domain. Using real-world tabular and text-based medical datasets, the study assesses these methods across six key metrics including identity, stability, and bias detection.

TL;DR

High-accuracy machine learning models in healthcare are often "black boxes" that clinicians struggle to trust. This paper benchmarks three industry-standard explanation tools—LIME, SHAP, and Anchors—using rigorous metrics like Stability and Identity. The verdict? There is no "perfect" explainer; SHAP leads in consistency and bias detection, while LIME and Anchors offer unique trade-offs in separability and rule-based clarity.

The Trust Gap in Medical AI

In medicine, a prediction without a "Why" is often a prediction without a use case. With the advent of GDPR, the "right to explanation" has moved from a moral preference to a legal requirement. However, while we have tools to explain models, we haven't reached a consensus on how to measure the quality of those explanations. Are they stable? Are they biased? Do they treat identical patients the same way?

Methodology: Benchmarking the Big Three

The researchers evaluated three distinct approaches to interpretability:

  1. LIME (Local Interpretable Model-agnostic Explanations): Approximates a complex model locally with a simple linear surrogate.
  2. SHAP (Shapley Additive Explanations): Uses coalitional game theory to distribute "credit" for a prediction among features.
  3. Anchors: Generates high-precision "if-then" rules that act as sufficient conditions for a prediction.

The study utilized real-world mortality and diabetes datasets (tabular) and drug reviews (text).

Evaluation Metrics

  • Identity: Ensure identical instances have identical explanations.
  • Stability: Ensure instances of the same class have similar explanations.
  • Separability: Ensure dissimilar instances result in different explanations.
  • Bias Detection: The ability of a human to spot "cheating" or "illogical" model behavior through the provided explanation.

Table of Results for Tabular Data

Key Insights: Strength and Weaknesses

1. The Identity Crisis of LIME

One of the most striking findings was LIME's 0% Identity score on tabular data. This occurs because LIME relies on random perturbations (sampling) around an instance. If you explain the same patient twice, LIME might give you two different answers. SHAP, conversely, achieved a perfect 100% Identity on tabular data.

2. Efficiency vs. Complexity

While Anchors provides very intuitive "rules," it is computationally expensive. In the diabetes dataset, Anchors took over 10 seconds per explanation, whereas SHAP and LIME handled it in approximately 0.2 seconds. For real-time clinical applications, this 50x delta is significant.

3. Spotting the "Smoker's Paradox" (Bias Detection)

The authors intentionally trained a biased mortality model where "Smoking" was incorrectly linked to lower mortality. They found that SHAP's visualization enabled 20 valid participants to detect this bias more effectively than LIME or Anchors.

Visual Interface for Bias Detection Figure 1: Comparison of interfaces used to detect bias in model explanations.

Conclusion and Recommendations

The research concludes that there is no "silver bullet."

  • Use SHAP if your priority is theoretical soundness, mathematical consistency (Identity), and bias detection.
  • Use LIME if you need high separability between different classes and can tolerate some sampling noise.
  • Use Anchors when the end-user (clinician) requires discrete, rule-based logic (e.g., "If Age > 60 and METS < 5, then Risk is High"), provided you have the computational overhead.

For the future of Healthcare AI, the message is clear: the choice of an "explainer" is as critical as the choice of the "classifier."

Critical Takeaway

The lack of a "clear winner" suggests that we should perhaps move toward hybrid interpretability pipelines where multiple methods are used to cross-verify clinical insights.

Find Similar Papers

Try Our Examples

  • Which recent papers propose new quantitative metrics for evaluating the consistency and faithfulness of local model-agnostic explanations beyond the ones used in this study?
  • How do the sampling-based approximations of SHAP specifically lead to the lower Identity scores observed in the text datasets compared to the tabular datasets in this research?
  • What are the current state-of-the-art methods for integrating local interpretability techniques into Clinical Decision Support Systems (CDSS) while maintaining GDPR compliance?
Contents
LIME, SHAP, or Anchors? Decoding the "Black Box" of Healthcare AI
1. TL;DR
2. The Trust Gap in Medical AI
3. Methodology: Benchmarking the Big Three
3.1. Evaluation Metrics
4. Key Insights: Strength and Weaknesses
4.1. 1. The Identity Crisis of LIME
4.2. 2. Efficiency vs. Complexity
4.3. 3. Spotting the "Smoker's Paradox" (Bias Detection)
5. Conclusion and Recommendations
5.1. Critical Takeaway