Boosting Emotion Detection: The Power of Weighted Ensembles in Thai Speech Recognition
Detecting Human Emotion via Speech Recognition by Using Ensemble Classification Model
This paper presents an Ensemble Classification Model for Speech Emotion Recognition (SER) specifically targeting Thai speech. By combining SVM, Neural Networks, and k-NN through a weighted majority voting scheme, the system achieves a peak accuracy of 70.69% in classifying four core emotions (anger, happy, natural, and sad) using MFCC and F0 features.
TL;DR
Recognizing human emotion from speech is a cornerstone of advanced Human-Computer Interaction (HCI). This research tackles the limitations of single-model classifiers by introducing a Weighted Majority Voting Ensemble that combines SVM, Neural Networks, and k-NN. Applied to a Thai emotional speech corpus, the method demonstrates that fusing MFCC (spectral) and F0 (prosodic) features yields a robust accuracy of 70.69%, significantly outperforming standalone models.
Context & Motivation
While Speech Emotion Recognition (SER) has seen various applications—from healthcare to automated tutoring—it remains a "hard" problem. Traditional models like GMMs or single SVMs often struggle with the high variance found in real-world audio, especially in tonal languages like Thai where background noise and music are present.
The authors' core insight is that diversity reduces error. By using an ensemble of classifiers with different theoretical backgrounds (geometric for SVM, instance-based for k-NN, and connectionist for NN), the system can cancel out individual model biases and produce a more reliable emotional "consensus."
Methodology: The Ensemble Architecture
The system follows a classic yet rigorous digital signal processing (DSP) pipeline:
- Pre-processing: High-pass filtering (pre-emphasis) to amplify high-frequency formants, followed by Hamming windowing to ensure signal continuity.
- Feature Fusion: The study compares several features including Energy, Zero-Crossing Rate (ZCR), and Pitch (F0). However, the "hero" feature is MFCC, which captures the power spectrum of the sound.
- Weighted Majority Vote: Unlike simple voting, this mechanism assigns a weight to each classifier based on its inherent accuracy .
- Formula:
- This ensures that the "expert" model for a specific feature set has a louder voice in the final decision.
Figure 1: The proposed SER system flow, from signal acquisition to ensemble decision.
Experimental Insights
The researchers utilized the Thai Emotional Speech Corpus, which includes realistic dialogue from drama shows. This is a higher bar than "clean" laboratory datasets.
Key Findings:
- MFCC is King: On its own, MFCC achieved ~68.97% accuracy in the ensemble, far higher than F0 (42.24%) or Energy (40.52%).
- Ensemble Synergy: The Weighted Majority Vote consistently edged out Bagging and single models. For example, using MFCC+F0, SVM reached 66.38%, while the Ensemble reached 70.69%.
- Emotion Sensitivity: The model is particularly good at spotting "Anger" (76.9%) and "Sadness" (73.9%), likely due to the distinct acoustic signatures these emotions leave on pitch and spectral energy.
Table 1: Comparison of different features and classification models. Note the peak at MFCC + F0 using Weighted Majority Vote.
Critical Analysis & Future Outlook
While the 70.69% accuracy is a solid step for Thai speech, the research also highlights the difficulty of "Neutral" and "Happy" emotions, which are often confused with one another in the confusion matrix.
Limitations: The current model still struggles with background music removal. In the "Thai drama" context, dramatic BGM can bleed into the spectral features of the speech.
Future Directions:
- Denoising: Implementing Deep Learning-based source separation (like Wave-U-Net) to isolate speech before feature extraction.
- Transformers: Moving from statistical features to self-supervised learning weights (like wav2vec 2.0) could likely push this accuracy beyond 80%.
Conclusion
This work reinforces the idea that in the messy reality of human speech, ensemble methods provide a necessary safety net against the instability of single models. For developers building HCI systems in tonal languages, the combination of MFCC and F0 remains the gold standard for feature selection.
