Decoding Emotions: Using XAI to Peer into the Black Box of Affective Computing

Explaining Machine Learning Models of Emotion Using the BIRAFFE Dataset

2020-01-01
Szymon Bobek, Magdalena M. Tragarz, Maciej Szelazek, Grzegorz J. Nalepa
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents a methodology for applying Explainable AI (XAI) to emotion detection models trained on the BIRAFFE dataset. It utilizes physiological signals (ECG and EDA) to classify emotions in a 2D Valence-Arousal space, achieving performance levels that highlight the inherent difficulty of affective computing.

TL;DR

Recognizing human emotions via wearable sensors is notoriously difficult due to signal noise and the subjectivity of feeling. This paper explores the BIRAFFE dataset, applying Explainable AI (XAI) techniques like SHAP, LIME, and Anchors to understand how models interpret ECG and Skin Conductance data. While the models themselves show that emotion "classification" is far from solved, the XAI tools provide a roadmap for better feature selection and model debugging.

The Problem: The "Black Box" of Bio-Signals

In Affective Computing (AfC), we often use the James-Lange theory: the idea that bodily changes are the emotion. If your heart races (ECG) and your palms sweat (EDA), you might be "Anxious" or "Excited."

However, building ML models for this faces two walls:

  1. Low Signal-to-Noise Ratio: Affordable wearables (like the BITalino used here) produce messy data compared to clinical tools.
  2. Model Opacity: When a model misclassifies "Sad" as "Calm," developers often don't know if the error stems from a bad feature (like BPM) or an inherent flaw in the model architecture.

Methodology: The BIRAFFE Protocol

The authors utilized the BIRAFFE (Bio-Reactions and Faces for Emotion-based Personalization) dataset, which recorded 206 participants.

  • Inputs: ECG (Heart rate metrics like BPM, SDNN, RMSSD) and EDA (Skin conductance peaks and tonic activity).
  • Stimuli: Standardized images (IAPS) and sounds (IADS).
  • Labels: A 2D Valence-Arousal space (the "Emospace") and a simpler 5-face scale ("Emoscale").

Model Architecture and Distribution Fig 1: The distribution of emotional labels across the Valence-Arousal quadrant.

The XAI Toolkit

To explain the predictions of their Random Forest and XGBoost models, they used:

  • SHAP: Assigns each feature an "importance" value based on cooperative game theory.
  • LIME: Creates a simplified, local linear model around a specific data point to see what's driving the decision.
  • Anchors: Finds the "if-then" rules (e.g., if BPM > 80 and EDA_peaks > 2, then 'Happy') that hold true for a large part of the data.

Results: A Reality Check for Affective ML

The results were humbling but insightful. For complex 4-class emotion tasks, most models hovered between "random" and "below average."

  • The "Sad" Gap: Models almost never correctly identified "Sad" emotions, often confusing them with "Calm."
  • Valence vs. Arousal: Models were much better at detecting Arousal (intensity of emotion) than Valence (positivity/negativity).
  • XGBoost Superiority: Among the tree-based models, XGBoost provided the most stable results and the highest local fidelity when analyzed by LIME.

Table of Extracted Features Table 1: The physiological features extracted from ECG and EDA signals.

Critical Insight: Why XAI Matters Here

The most valuable part of this study isn't the accuracy of the models—it's the local fidelity analysis. By using Anchors, the authors could see that certain decisions were based almost entirely on a single feature, like amplitude_avg. If that feature is noisy, the whole model collapses.

The Anchor mechanism showed that for some successful predictions, the "coverage" (the percentage of similar cases the rule applies to) was as high as 38%, giving researchers a clear target for feature engineering.

Conclusion & Future Outlook

This work demonstrates that while we aren't yet at a point where a smartwatch can perfectly read your feelings, Explainable AI provides the diagnostic tools needed to get there.

Key Takeaways for Researchers:

  • Don't trust global accuracy: A model might have high accuracy just by guessing the majority class (e.g., "Positive Valence").
  • Use Anchors for Stability: If your model's explanation changes wildly between two similar data points, your model hasn't "learned" the emotion; it's just overfitting.
  • Next Step - Personalization: The authors suggest that the future of AfC lies in context-dependent and personalized models, where XAI helps tailor the algorithm to an individual's unique physiological "fingerprint."

Perspectives

The BIRAFFE dataset remains a vital open-source resource for the community. You can access it at Zenodo.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize the BIRAFFE dataset or similar multimodal affective datasets to improve emotion classification accuracy using Deep Learning.
  • What are the current state-of-the-art methods for "context-aware" physiological emotion recognition that address the limitations mentioned in the James-Lange approach?
  • Find papers that apply Anchor-based explanations to other biometric time-series data, such as EEG or gait analysis, to compare feature stability.
Contents
Decoding Emotions: Using XAI to Peer into the Black Box of Affective Computing
1. TL;DR
2. The Problem: The "Black Box" of Bio-Signals
3. Methodology: The BIRAFFE Protocol
3.1. The XAI Toolkit
4. Results: A Reality Check for Affective ML
5. Critical Insight: Why XAI Matters Here
6. Conclusion & Future Outlook
7. Perspectives