Bridging the Emotional Gap: Scaling Speech Recognition with Twitter Intelligence
Language Model Adaptation for Emotional Speech Recognition using Tweet data
This paper proposes a language model (LM) adaptation method using a large-scale Twitter dataset (25.86M words) to improve Japanese Emotional Speech Recognition. By leveraging the colloquial and emotional nature of tweets, the authors significantly reduce the Word Error Rate (WER) on the JTES and OGVC corpora.
TL;DR
Recognizing emotional speech has long been a "stumbling block" for ASR systems due to the vast difference between formal training data and real-world outbursts. This paper presents a breakthrough by using 25.86 million words of Twitter data to adapt Language Models, cutting Word Error Rates (WER) from 36.11% down to 17.77%.
Background: Why "Normal" ASR Fails at Emotion
Most Japanese ASR systems are trained on the Corpus of Spontaneous Japanese (CSJ)—essentially a collection of academic lectures. While great for formal settings, it is "emotionally sterile." When a speaker gets angry, joyful, or sad, two things happen:
- Acoustic Shift: Pitch, intensity, and duration change drastically.
- Linguistic Shift: Speakers use "colloquialisms," drop particles (like ga or wo), and emphasize specific Japanese phonemes like the geminate stop /Q/ (baQkari).
Previous studies failed because they tried to adapt models using tiny datasets (under 2,000 sentences). The authors’ core insight is that volume beats precision: a massive amount of unlabeled, informal text (Tweets) is more valuable than a handful of perfectly labeled emotional sentences.
Methodology: The Two-Pronged Adaptation
1. Large-Scale Twitter LM Adaptation
The authors used the Twitter API to collect 51 days of Japanese tweets. After stripping URLs and hashtags, they applied a Bigram Perplexity filter using the CSJ model to ensure the selected sentences were linguistically coherent while retaining their "human" emotional flavor.
The adaptation uses a Mixed N-gram approach, which mathematically balances the baseline (formal) counts with the new (tweet) counts: This allows the model to "remember" standard Japanese grammar while "learning" the emotional shortcuts used on social media.
2. Acoustic Model (AM) Adaptation
The system utilizes a DNN-HMM architecture. To handle the acoustic variance, the authors performed supervised backpropagation on the JTES (Japanese Twitter-based emotional speech) corpus. A key innovation here is Output Probability Compensation, which prevents common states (like silence) from overwhelming the model's predictions during emotional peaks.
Table: Test set perplexity significantly improves (from ~900 to ~224) when switching to large-scale tweet adaptation.
Experimental Results: Quantitative Dominance
The results on the JTES corpus were striking across all emotion types (Anger, Joy, Neu, Sad):
| Method | Average WER (%) |
|---|---|
| Baseline (CSJ) | 36.11 |
| LM Adaptation (Large-Scale) | 25.68 |
| Combined AM + LM Adaptation | 17.77 |
Table: Results show that "Sadness" and "Anger" saw the most dramatic improvements, likely due to the higher frequency of colloquial markers in those states.
Qualitative Win: Handling the "Nuance"
The paper highlights specific cases where the model succeeded:
- Particle Dropping: Correctly identifying "zikan aru toki" instead of the formal "zikan ga aru toki."
- Emphasis: Capturing the emphatic stop /Q/ in "baQkari."
- Morphofogical Accuracy: Avoiding errors in auxiliary verbs like "na no ni."
Deep Insight & Conclusion
The most significant takeaway is Versatility. Even when tested on the OGVC (Online Gaming Voice Chat) corpus—a completely different environment from Twitter—the model still outperformed the baseline. This suggests that the "emotional language" of Twitter is a generic proxy for "informal human interaction."
Limitations: The model still struggles with "fillers" (long vowels) and extreme emotional intensity (Intensity Level 3). The authors suggest that future work should focus on Emotion-Dependent AMs and upgrading from N-grams to Transformer-based Neural LMs.
Takeaway for Practitioners: If you are building ASR for real-world interactions, don't just look for speech data. A massive crawl of informal text from the target culture may be the cheapest and most effective "performance booster" for your language model.
