DEMoS: Bridging the Gap Between Acted and Natural Emotional Speech in Italian

DEMoS – An Italian Emotional Speech Corpus Elicitation methods, machine learning, and perception

2020-01-01
Emilia Parada-Cabaleiro, Giovanni, Costantini, A. Batliner, Maximilian Schmitt, Björn Schuller
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces DEMoS (Database of Elicited Mood in Speech), a large-scale Italian emotional speech corpus containing ~9k samples from 68 speakers. It employs multi-modal mood induction procedures (MIPs) and evaluates the corpus using Support Vector Machines (SVM) and perceptual tests, achieving a comprehensive benchmark for Italian speech emotion recognition (SER).

TL;DR

The DEMoS project addresses the scarcity of Italian emotional speech data by creating a corpus of 9,365 samples elicited through psychological induction rather than acting. By testing various Machine Learning (ML) strategies, the study reveals a counter-intuitive truth: data volume and diversity (even if "noisy" or non-typical) are more important for model performance than the "purity" of emotional prototypes.

Background: The Authenticity Paradox

In Speech Emotion Recognition (SER), researchers usually face a dilemma. Do you use acted speech (easy to collect, clear labels, but fake) or natural speech (real, but noisy and legally complex)?

The authors position DEMoS as a middle ground. By using Mood Induction Procedures (MIPs)—like making a subject recall a sad memory while listening to minor-key music—they capture speech that is both acoustically clean and psychologically grounded.

Methodology: The Science of Feeling

The researchers didn't just ask people to "be angry." They designed three distinct elicitation methods:

  1. Method A (Happiness, Sadness, Guilt, Surprise): Music + Autobiographical Recall + Empathy texts.
  2. Method B (Anger, Fear): Film clips + Self-statements.
  3. Method C (Disgust): Subjective imagery (Medical/Nature pictures).

The Prototypicality Pipeline

To find the "perfect" examples of emotion, the authors used a multi-stage filter:

  • Alexithymia Test (SAR): Discarded speakers who can't accurately identify emotions.
  • Self-Assessment: Did the speaker feel what we tried to induce?
  • External Assessment: Do experts agree it sounds like that emotion?

Model Architecture and Elicitation Workflow

Machine Learning Insights: Quality vs. Quantity

Using the openSMILE ComParE feature set (6,373 acoustic features) and SVMs, the team ran 68 speaker-dependent experiments.

Key ML Findings:

  • Sadness and Anger are "Low-Hanging Fruit": These achieved the highest recalls.
  • The Power of Variability: In the "Inter-group" test, training on "non-prototypical" samples (the "messy" data) actually made the models more robust.
  • Quantity Wins: Models trained on the full dataset (9k samples) significantly outperformed those trained only on the highly verified "prototypical" subset (1.5k samples).

Experimental Results Comparison

Acoustic and Perceptual Reality

The acoustic analysis of F0 range and Energy confirmed that "selected" (prototypical) samples have very distinct signatures, whereas "non-selected" samples blur together.

Interestingly, a Perceptual Study with 86 native listeners showed that humans often confuse Guilt with Sadness and Surprise with Happiness. This highlights the limitations of the "Big 6" categorical model—real emotions are often secondary or ambiguous, and the 2D Arousal-Valence space isn't always enough to distinguish them.

Acoustic Feature Comparison

Critical Analysis & Conclusion

Takeaway: The DEMoS corpus proves that elicited speech is a viable, high-quality alternative to both acting and "wild" data. For tech developers, the message is clear: Don't over-curate your training data. The "less typical" samples provide the inductive bias necessary for the model to handle the ambiguity of real-world Italian speech.

Limitations:

  • Eliciting Fear remains notoriously difficult in a lab setting.
  • The use of read-aloud texts (Empathy MIP) creates a "reading style" bias that might differ from spontaneous conversation.

Future Work: Expanding DEMoS beyond categorical labels into continuous dimensional ratings and exploring how deep-learning-based feature extraction (like Wav2Vec) handles the nuances of Italian prosody.

Find Similar Papers

Try Our Examples

  • Search for recent Italian emotional speech corpora published after 2019 that utilize deep learning architectures like Wav2Vec 2.0 or HuBERT.
  • Which original papers established the ComParE feature set in the openSMILE toolkit, and how do these features compare to modern self-supervised learning representations for emotion recognition?
  • Investigate how mood induction procedures (MIPs) are currently being adapted for cross-cultural emotional speech collection in under-resourced languages.
Contents
DEMoS: Bridging the Gap Between Acted and Natural Emotional Speech in Italian
1. TL;DR
2. Background: The Authenticity Paradox
3. Methodology: The Science of Feeling
3.1. The Prototypicality Pipeline
4. Machine Learning Insights: Quality vs. Quantity
4.1. Key ML Findings:
5. Acoustic and Perceptual Reality
6. Critical Analysis & Conclusion