Emo-Soundscapes: Decoding the Emotional Pulse of Environmental Sound

Emo-soundscapes: A dataset for soundscape emotion recognition

2017-10-01
Jianyu Fan, Miles Thorogood, Philippe Pasquier
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Emo-Soundscapes, a public dataset specifically designed for Soundscape Emotion Recognition (SER). It utilizes a relative ranking-based crowdsourcing methodology to annotate 1,213 audio clips across valence and arousal dimensions, establishing a new SOTA baseline for predicting perceived emotions in environmental audio.

TL;DR

Recognizing the emotion conveyed by a "soundscape"—the acoustic environment we inhabit—is critical for urban design and human-computer interaction. This paper presents Emo-Soundscapes, the first major open-access dataset for Soundscape Emotion Recognition (SER). By utilizing a clever pairwise ranking system instead of traditional ratings, the authors provide 1,213 high-quality annotated clips and a robust SVR baseline that pushes the R² performance boundary for both valence and arousal.

Problem & Motivation: The Subjectivity Trap

Understanding why a park feels "calm" while a factory feels "agitated" is the core of SER. However, the field has been plagued by two major issues:

  1. Data Scarcity: Previous datasets were often private or extremely limited in diversity.
  2. The Rating Flaw: Asking an annotator to rate "Valence" on a scale of 1-10 is notoriously unreliable. One person's "7" is another's "5," especially across different cultures.

The authors argue that human beings are much better at comparing two things ("Is clip A more pleasant than clip B?") than assigning absolute values. This insight led to a pivot toward a rank-based ground truth.

Methodology: Taxonomy and Quick-Sort Crowdsourcing

1. Dataset Composition

The authors selected 600 clips based on Schafer’s soundscape taxonomy, covering categories like:

  • Natural sounds (birds, rain)
  • Human sounds (laughter)
  • Sounds as indicators (bells)

They then created 613 "mixed" recordings (e.g., mixing wind with mechanical noise) to investigate how overlapping sound sources shift emotional perception.

2. The Ranking Interface

To efficiently rank 1,213 clips, they used a Quick Sort logic. In each iteration, clips were compared against a "pivot." This reduced the complexity of comparisons needed to reach a global order of valence and arousal.

Interface for Crowdsourcing Workers Figure 1: The annotation interface using the Self-Assessment Manikin (SAM) to guide workers in ranking valence and arousal.

Experiments & Results: Setting the Baseline

The authors extracted 122 acoustic features (MFCCs, spectral flux, etc.) and narrowed them down to a 39-dimension vector for training a Support Vector Regression (SVR) model.

Key Breakthroughs:

  • Arousal Prediction: Achieved an R² of 0.855, indicating that energetic/eventful sounds are highly predictable via acoustic features.
  • Valence Prediction: Achieved an R² of 0.629, which is significantly higher than previous studies (typically <0.57).

Table of Arousal and Valence Metrics Figure 2: Baseline performance of SVR on Protocol A (Shuffle) for Emo-Soundscapes.

Critical Analysis & Conclusion

Emo-Soundscapes fills a vital gap by providing a reproducible benchmark. The shift to ordinal (ranking) data is a masterstroke for handling crowdsourced subjectivity.

Limitations: The current rankings are relative. A clip ranked "1st" is the most pleasant in this set, but we don't know its absolute intensity. The authors are already planning follows-up to correlate these rankings with absolute emotional scales.

Future Outlook: With the rise of "Smart Cities," SER models trained on this dataset could automatically monitor urban well-being, helping planners design environments that actively reduce stress and improve public acoustic health.

Find Similar Papers

Try Our Examples

  • Find recent papers on Soundscape Emotion Recognition that utilize Deep Learning or Transformer-based architectures instead of traditional SVR.
  • Which paper originally proposed the "Ranking is better than Rating" hypothesis for affective computing, and how does Emo-Soundscapes expand on that methodology?
  • Explore research that applies the Emo-Soundscapes dataset or similar environmental audio data to urban planning and noise pollution mitigation strategies.
Contents
Emo-Soundscapes: Decoding the Emotional Pulse of Environmental Sound
1. TL;DR
2. Problem & Motivation: The Subjectivity Trap
3. Methodology: Taxonomy and Quick-Sort Crowdsourcing
3.1. 1. Dataset Composition
3.2. 2. The Ranking Interface
4. Experiments & Results: Setting the Baseline
4.1. Key Breakthroughs:
5. Critical Analysis & Conclusion