Boosting Emotion Detection: The Power of Weighted Ensembles in Thai Speech Recognition

Detecting Human Emotion via Speech Recognition by Using Ensemble Classification Model

2018-01-01
Sathit Prasomphan, Surinee Doungwichain
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents an Ensemble Classification Model for Speech Emotion Recognition (SER) specifically targeting Thai speech. By combining SVM, Neural Networks, and k-NN through a weighted majority voting scheme, the system achieves a peak accuracy of 70.69% in classifying four core emotions (anger, happy, natural, and sad) using MFCC and F0 features.

TL;DR

Recognizing human emotion from speech is a cornerstone of advanced Human-Computer Interaction (HCI). This research tackles the limitations of single-model classifiers by introducing a Weighted Majority Voting Ensemble that combines SVM, Neural Networks, and k-NN. Applied to a Thai emotional speech corpus, the method demonstrates that fusing MFCC (spectral) and F0 (prosodic) features yields a robust accuracy of 70.69%, significantly outperforming standalone models.

Context & Motivation

While Speech Emotion Recognition (SER) has seen various applications—from healthcare to automated tutoring—it remains a "hard" problem. Traditional models like GMMs or single SVMs often struggle with the high variance found in real-world audio, especially in tonal languages like Thai where background noise and music are present.

The authors' core insight is that diversity reduces error. By using an ensemble of classifiers with different theoretical backgrounds (geometric for SVM, instance-based for k-NN, and connectionist for NN), the system can cancel out individual model biases and produce a more reliable emotional "consensus."

Methodology: The Ensemble Architecture

The system follows a classic yet rigorous digital signal processing (DSP) pipeline:

  1. Pre-processing: High-pass filtering (pre-emphasis) to amplify high-frequency formants, followed by Hamming windowing to ensure signal continuity.
  2. Feature Fusion: The study compares several features including Energy, Zero-Crossing Rate (ZCR), and Pitch (F0). However, the "hero" feature is MFCC, which captures the power spectrum of the sound.
  3. Weighted Majority Vote: Unlike simple voting, this mechanism assigns a weight to each classifier based on its inherent accuracy .
    • Formula:
    • This ensures that the "expert" model for a specific feature set has a louder voice in the final decision.

Overall Architecture Figure 1: The proposed SER system flow, from signal acquisition to ensemble decision.

Experimental Insights

The researchers utilized the Thai Emotional Speech Corpus, which includes realistic dialogue from drama shows. This is a higher bar than "clean" laboratory datasets.

Key Findings:

  • MFCC is King: On its own, MFCC achieved ~68.97% accuracy in the ensemble, far higher than F0 (42.24%) or Energy (40.52%).
  • Ensemble Synergy: The Weighted Majority Vote consistently edged out Bagging and single models. For example, using MFCC+F0, SVM reached 66.38%, while the Ensemble reached 70.69%.
  • Emotion Sensitivity: The model is particularly good at spotting "Anger" (76.9%) and "Sadness" (73.9%), likely due to the distinct acoustic signatures these emotions leave on pitch and spectral energy.

Performance Comparison Table 1: Comparison of different features and classification models. Note the peak at MFCC + F0 using Weighted Majority Vote.

Critical Analysis & Future Outlook

While the 70.69% accuracy is a solid step for Thai speech, the research also highlights the difficulty of "Neutral" and "Happy" emotions, which are often confused with one another in the confusion matrix.

Limitations: The current model still struggles with background music removal. In the "Thai drama" context, dramatic BGM can bleed into the spectral features of the speech.

Future Directions:

  • Denoising: Implementing Deep Learning-based source separation (like Wave-U-Net) to isolate speech before feature extraction.
  • Transformers: Moving from statistical features to self-supervised learning weights (like wav2vec 2.0) could likely push this accuracy beyond 80%.

Conclusion

This work reinforces the idea that in the messy reality of human speech, ensemble methods provide a necessary safety net against the instability of single models. For developers building HCI systems in tonal languages, the combination of MFCC and F0 remains the gold standard for feature selection.

Find Similar Papers

Try Our Examples

  • Search for recent papers on Speech Emotion Recognition specifically for tonal languages using deep learning ensemble methods.
  • Which study first established Mel-Frequency Cepstral Coefficients (MFCC) as a superior feature to prosodic features for emotion detection, and how has this evolved with Transformer-based architectures?
  • Explore research that applies weighted majority voting ensembles to emotion recognition in multi-modal contexts, such as combining audio with facial expression analysis.
Contents
Boosting Emotion Detection: The Power of Weighted Ensembles in Thai Speech Recognition
1. TL;DR
2. Context & Motivation
3. Methodology: The Ensemble Architecture
4. Experimental Insights
4.1. Key Findings:
5. Critical Analysis & Future Outlook
6. Conclusion