Overcoming the "Day-Effect": How Multi-Day Training Stabilizes EEG Emotion Recognition
12187_Improve the generalization of emotional classifiers across time by using training samples from different days.
This paper presents a study on enhancing the temporal generalization of EEG-based emotion recognition using a multi-day training strategy. The core method involves pooling EEG samples from multiple experimental sessions (up to 4 days) to train a Support Vector Machine (SVM), achieving a peak average accuracy of 73.0% for three emotional states (positive, neutral, negative).
TL;DR
Recognizing emotions from EEG signals is notoriously difficult because our brain patterns change daily. This study demonstrates that by simply including training data from multiple different days (up to 4 days), we can significantly improve a classifier's ability to recognize emotions on a future, unseen day. Using an SVM-based approach, the researchers achieved a 10% accuracy boost, reaching 73.0% for a three-class emotion task.
Context: The Challenge of Temporal Non-Stationarity
In the world of Brain-Computer Interaction (BCI), we often encounter a frustrating phenomenon: a model that works perfectly today might be useless tomorrow. This is known as the "day-effect." Factors such as hormone levels, sleep quality, and electrode impedance cause the "baseline" EEG signal to shift.
Most prior works focus on intra-session classification (training and testing on the same day). This paper tackles the more realistic "cross-day" challenge, asking: How can we build a model that isn't fooled by the passage of time?
Methodology: Diversity as a Filter
The authors recruited 8 subjects for a longitudinal study involving 5 sessions over one month. They used movie clips to elicit Positive, Neutral, and Negative states.
The core hypothesis is elegant: if a classifier only sees data from one day, it will overfit to the specific "noise" of that day. If it sees data from 4 different days, the only consistent signal remaining across those sessions must be the emotional state itself.
The Pipeline:
- Feature Extraction: Power Spectral Density (PSD) across 6 frequency bands (delta to gamma).
- Classification: Support Vector Machine (SVM) with a Gaussian RBF kernel.
- Experimental Conditions: Testing the model's performance as the training set grows from 1 day of data to 4 days of data.
Figure 1: The training setup involved alternating between training on N days and testing on the remaining (5-N) days to ensure unbiased results.
Why It Works: Feature Stability
The researchers used SVM-RFE (Recursive Feature Elimination) to rank features. Their findings confirmed their intuition:
- Stable Features: Features selected consistently across multiple days showed distinct, repeatable gaps between Positive, Neutral, and Negative states regardless of when they were recorded.
- Unstable Features: Discarded features showed "reversals"—for example, showing high power for "Positive" on Day 2 but low power for "Positive" on Day 3.
Figure 2: The clear ascending trend shows that as the variety of training days increases, the model's accuracy on unseen future days improves consistently for all subjects.
Results and Insights
- The 10% Jump: Moving from 1-day training to 4-day training boosted accuracy from 64.9% to 73.0%.
- Feature Selection is Key: The model successfully learned to "reject" time-relevant features (noise) and "select" emotion-relevant ones.
- Subject Variability: While all subjects improved, the peak accuracy reached 81.2% for some individuals, suggesting that some brains have more "stable" emotional signatures than others.
Figure 3: Stable features (top) vs. Unstable features (bottom). Notice how the stable features maintain their relative order (P > N > Ne) across all sessions.
Evaluation and Future Work
This study provides a practical "brute force" solution to the day-effect: data diversity. However, it also highlights a limitation—collecting data over multiple days is time-consuming for the user.
The next frontier in this research will likely involve Domain Adaptation (DA) or Generative Adversarial Networks (GANs) to "synthesize" the day-effect variations, allowing us to achieve this high generalization without requiring users to participate in week-long calibration sessions.
Takeaway for Practitioners
If you are building an EEG-based product, do not rely on a single calibration session. Longitudinal data is not just "more data"; it is "different information" that allows your model to distinguish between a user's mood and their biological baseline.
