ELM for Emotion Recognition: Speed and Accuracy in the Wild
Combining modality-specific extreme learning machines for emotion recognition in the wild
Deeply engaged in the ICMI 2014 Emotion Recognition in the Wild (EmotiW) Challenge, this paper proposes a multimodal system utilizing Extreme Learning Machines (ELM) for audio-visual emotion classification. The authors demonstrate that Kernel ELMs efficiently model modality-specific features, achieving competitive accuracies of 44.23% on the challenge test set through score-level fusion.
TL;DR
Recognizing human emotions in real-world movie clips is notoriously difficult due to environmental noise and occlusion. This paper leverages Extreme Learning Machines (ELM)—a fast learning paradigm—to process audio-visual data. By focusing on the inner facial regions and employing Kernel ELM, the authors achieved a 44.23% accuracy on the AFEW dataset, proving that smart feature selection and fast solvers can compete with more complex models.
Context: Moving from the Lab to the "Wild"
Prior to the EmotiW challenges, most emotion recognition research focused on "clean" data (e.g., actors in studios). However, real life is messy. The Acted Facial Expression in the Wild (AFEW) dataset uses movie clips to simulate this complexity. The researchers identified that standard SVMs and DNNs, while powerful, are often slow to iterate. They turned to ELMs to explore more hypotheses—like which part of the face matters most—using moderate hardware.
Methodology: The Power of Extreme Learning Machines
The core of this work is the Extreme Learning Machine (ELM). Unlike traditional back-propagation networks, ELM randomly generates or uses PCA for the input layer weights and solves the output weights using a simple Least Squares solution.
1. Visual Strategy: Why the "Inner Face"?
The authors hypothesized that the full face contains "cluttered" information—hair, backgrounds, or occlusions at the edges. By dividing the face into 16 regions (4x4) and selecting only the "inner face" (eyes and mouth), they reduced dimensionality while focusing on regions where emotional actions (Action Units) occur.
Figure 1: Comparison of different facial region groupings tested for robustness.
2. Acoustic Strategy: Quality over Quantity
The team tried augmenting the data with other emotional corpora (like EMODB). Surprisingly, this didn't help much because those corpora were recorded in "laboratory" conditions, which are too different from movie data. Instead, using Kernel ELMs and CCA-based feature selection proved more effective.
Experimental Insights
The results demonstrated that Kernel ELMs (RBF) consistently outperformed basic ELMs. A crucial discovery was that PCA-initialized input weights provided a substantial boost over random initialization, suggesting that even in "random" networks, a data-driven starting point helps significantly.
Table 1: Performance comparison between Basic ELM, PCA-ELM, and Kernel versions.
In the final test set, the multimodal fusion (combining audio and video scores) reached 44.23%, a significant jump from the individual modalities.
Critical Analysis & Conclusion
Takeaway
This paper serves as a reminder that feature engineering and fast learning algorithms still have a place in the era of Deep Learning. By focusing on the "inner face," the authors built a system that is naturally more resilient to the "wild" conditions of the dataset.
Limitations
The system struggled with specific classes like Disgust and Surprise in the audio modality (as seen in the confusion matrix). This indicates that the current acoustic features (MFCCs, F0) might not capture the nuances of these specific emotions as well as visual features do.
Future Outlook
While this work used LBP-TOP features (state-of-the-art in 2014), the logic of using ELM as a fast backend for multimodal fusion remains highly relevant today for edge computing and real-time emotion sensing where GPU resources are limited.
Figure 2: Final Fusion System Confusion Matrix showing strengths in Anger and Neutral classes.
