Bridging the Gap: Evaluating XAI Through the Eyes of the Crowd

Crowdsourcing Evaluation of Saliency-Based XAI Methods

2021-01-01
Xiaotian Lu, Arseny Tolmachev, Tatsuya Yamamoto, Koh Takeuchi, Seiji Okajima, Tomoyoshi Takebayashi, Koji Maruhashi, Hisashi Kashima
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a novel crowd-based evaluation scheme for saliency-based XAI methods, inspired by the "Peek-a-boom" human computation game. It quantitatively assesses how effectively different saliency maps guide humans toward correct image classification at varying pixel exposure rates, identifying Grad-CAM as the most human-interpretable method among those tested.

TL;DR

Is a "good" AI explanation one that a machine understands, or one a human finds useful? This paper argues for the latter, proposing a crowdsourced evaluation framework based on the game Peek-a-boom. By testing four major XAI methods, the authors discovered that popular automated benchmarks like ROAR don't always align with human intuition, and that Grad-CAM—despite being lower resolution—is the most effective at helping humans understand image content.

The "Ground Truth" Problem in XAI

The field of Explainable AI (XAI) is booming, yet we lack a "gold standard" for measuring success. Most researchers use automated metrics: they delete pixels marked as "important" by an algorithm and see how fast the AI's accuracy drops.

But there is a fundamental flaw in this logic: Machines and humans see differently. A model might classify a plane based on the "sky" pixels, a statistically valid but humanly confusing explanation. As the authors point out, interpretability is a human-centric property, yet we rarely involve actual humans in the evaluation loop due to cost and complexity.

Methodology: The Peek-a-boom Strategy

To solve this, the researchers adapted a human computation game called Peek-a-boom.

How it Works:

  1. Ranking: An XAI method (like SmoothGrad) ranks pixels by importance.
  2. Incremental Revelation: A crowd worker is shown just 5% of the "most important" pixels.
  3. The Guessing Game: The worker tries to classify the image. If they can't, more pixels are revealed (e.g., 10%, 15%) until they answer correctly.
  4. Metric: The target is a high Accuracy-Exposure Curve. If a human can identify a "Golden Retriever" with only 10% of the image visible, the XAI method is highly effective at capturing human-relevant features.

Peek-a-boom Interface Figure: The proposed crowdsourcing interface where pixels are revealed based on saliency rankings.

Key Findings: Grad-CAM Wins, But Why?

The study compared four heavyweights: Vanilla Gradient, SmoothGrad, Guided-Backpropagation, and Grad-CAM.

  • The Winner: Grad-CAM consistently outperformed the others in human trials.
  • The Intuition: Fine-grained methods (like Guided-BP) highlight sharp edges and outlines. While mathematically precise, they are often "noisy" to the human eye. Grad-CAM focuses on blobs/regions. For a human, seeing a blurry but localized "head" of an animal is more useful than seeing a scattered set of high-contrast edges.

Comparison of Saliency Maps Figure: Visual comparison showing how different methods highlight features in the same image.

Automated vs. Human: The Disconnect

The most striking part of the paper is the comparison between human results and automated metrics (ROAR, KAR, ROAE, KAE).

The authors found that ROAR (Remove and Retrain)—often considered the industry standard—yielded rankings that were significantly different from human judgment. Instead, KAE (Keep and Evaluate) showed the highest correlation with how humans performed. This suggests that the XAI community might be over-relying on evaluation metrics that don't actually reflect human interpretability.

Results Summary Table Table: Quantitative comparison (AUC) showing Grad-CAM's dominance in the "Crowd" column.

Critical Insight & Future Outlook

This work serves as a reality check for XAI developers. Performance on a benchmark does not guarantee "trust" from a human user.

Takeaways for Practitioners:

  1. Don't trust ROAR blindly: If your goal is a human-facing application (e.g., medical diagnosis assistance), prioritize human-in-the-loop testing.
  2. Stability of Crowds: The study found that even with varying worker skill levels, the aggregate "wisdom of the crowd" remains stable, making this a viable evaluation path.
  3. Region over Edges: For visual explanations, coarse region-based highlights are often superior to high-frequency gradient maps for human communication.

Limitations: The study focuses on image classification. Whether these findings translate to complex reasoning tasks or non-visual data remains an open question for future research.

Find Similar Papers

Try Our Examples

  • Find recent papers that propose alternative human-in-the-loop evaluation frameworks for XAI beyond saliency maps, specifically for NLP or tabular data.
  • What is the theoretical origin of the 'Remove and Retrain' (ROAR) framework, and how have subsequent studies addressed its high computational overhead?
  • Search for studies investigating whether Grad-CAM's superior human interpretability holds true in high-stakes domains like medical imaging or autonomous driving.
Contents
Bridging the Gap: Evaluating XAI Through the Eyes of the Crowd
1. TL;DR
2. The "Ground Truth" Problem in XAI
3. Methodology: The Peek-a-boom Strategy
3.1. How it Works:
4. Key Findings: Grad-CAM Wins, But Why?
5. Automated vs. Human: The Disconnect
6. Critical Insight & Future Outlook