CopyPaste: Leveraging Perceptual Salience for Robust Speech Emotion Recognition

CopyPaste: An Augmentation Method for Speech Emotion Recognition

2021-05-13
Raghavendra Pappagari, Jesús Villalba, Piotr Żelasko, Laureano Moro-Velázquez, Najim Dehak
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces CopyPaste, a perceptually motivated data augmentation method for Speech Emotion Recognition (SER). By concatenating emotional and neutral speech segments, the authors leverage the psychological insight that non-neutral emotions dominate human perception of an utterance, consistently achieving SOTA improvements across MSP-Podcast, Crema-D, and IEMOCAP datasets.

TL;DR

Data scarcity is the "Achilles' heel" of Speech Emotion Recognition (SER). This paper presents CopyPaste, a novel augmentation technique that creates new training samples by concatenating emotional segments with neutral or similar emotional utterances. This approach is rooted in the human perceptual bias where short emotional bursts define the overall perceived emotion of a longer clip. The results show consistent performance boosts across multiple benchmarks, especially when combined with traditional noise augmentation and transfer learning.

Background: Why SER is a Data-Starved Domain

Unlike Automatic Speech Recognition (ASR), where thousands of hours of data are readily available, SER relies on datasets that are notoriously small and expensive to produce. Emotions are ephemeral and subjective; annotating a few hours of speech requires multiple experts to achieve consensus.

The authors identify a specific technical gap: in a 10-second clip where a speaker is angry for 3 seconds and neutral for 7, a human will label the whole clip as "Angry." However, standard machine learning models often struggle as the "Neutral" statistics dominate the utterance-level representation.

Methodology: The CopyPaste Logic

The authors utilize a ResNet-34 architecture equipped with a Multi-head Attention Pooling mechanism. This architecture is designed to map frame-level features (MFCCs) into a fixed-length utterance embedding.

The Three Schemes

  1. Neutral CopyPaste (N-CP): Concatenating an emotional utterance (E) with a neutral one. Label = E. Insight: Forces the model to ignore neutral "noise" and focus on the emotional signal.
  2. Same Emotion CopyPaste (SE-CP): Concatenating two different utterances of the same emotion. Label = E. Insight: Increases phonetic and speaker diversity within a single emotional class.
  3. N+SE-CP: A hybrid approach using both methods.

Model Architecture

Table 1: The ResNet-34 architecture used for both independent training and x-vector transfer learning.

Experiments and Results

The researchers tested their hypothesis on three major datasets: MSP-Podcast (naturalistic), Crema-D (acted), and IEMOCAP (improvised/scripted).

Key Findings

  • Transfer Learning Synergy: While pre-training on speaker recognition (x-vectors) provides a strong baseline, CopyPaste consistently adds marginal value, pushing the F1-scores higher across the board.
  • Noise Robustness: Interestingly, CopyPaste outperformed standard noise augmentation in most clean-test scenarios. When the test environment was corrupted with 0dB/10dB noise, the best results were obtained by using both noise and CopyPaste during training.
  • Per-Class Improvement: Evaluation on Crema-D showed that the Neutral CopyPaste (N-CP) scheme improved scores for all classes (Angry, Happy, Sad, Neutral), suggesting it doesn't just help the model find "Anger"—it helps the model distinguish all emotions from the "Neutral" baseline.

Performance Comparison Table 2: Significant improvements observed when applying CopyPaste schemes to the pre-trained ResNet model.

Critical Analysis & Conclusion

Takeaway

The success of CopyPaste lies in its Inductive Bias. By artificially creating utterances where the "meaningful" content (emotion) is sparse compared to the total length, the authors train the Attention mechanism of the ResNet to be more selective. This mimics the "Peak-End Rule" of human psychology, where we judge experiences based on their most intense points rather than the average.

Limitations

While effective, the study primarily uses 4-second segments for concatenation. In real-world "in-the-wild" scenarios, emotional bursts might be much shorter (sub-second). Furthermore, the risk of biasing the model against the "Neutral" class—since neutral segments are used as "background" for all other classes—requires careful batch-level probability tuning as mentioned by the authors.

Future Outlook

This "Perceptual CopyPaste" logic isn't limited to SER. It holds potential for any utterance-level task where the target label is stay-invariant under concatenation with "background" noise, such as Language Identification or Speaker Verification in multi-speaker environments.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply temporal segment concatenation or "mixup" strategies specifically for speech emotion recognition tasks.
  • Which original research established that human listeners classify sequences of neutral and emotional speech based on the non-neutral segment, and how has this been validated in recent psychoacoustic studies?
  • Explore studies that evaluate the robustness of x-vector based transfer learning for other non-speaker tasks like involuntary physiological signal detection or age estimation.
Contents
CopyPaste: Leveraging Perceptual Salience for Robust Speech Emotion Recognition
1. TL;DR
2. Background: Why SER is a Data-Starved Domain
3. Methodology: The CopyPaste Logic
3.1. The Three Schemes
4. Experiments and Results
4.1. Key Findings
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook