ELM for Emotion Recognition: Speed and Accuracy in the Wild

Combining modality-specific extreme learning machines for emotion recognition in the wild

2015-04-29
Heysem Kaya, Albert Ali Salah
Summary
Problem
Method
Results
Takeaways
Abstract

Deeply engaged in the ICMI 2014 Emotion Recognition in the Wild (EmotiW) Challenge, this paper proposes a multimodal system utilizing Extreme Learning Machines (ELM) for audio-visual emotion classification. The authors demonstrate that Kernel ELMs efficiently model modality-specific features, achieving competitive accuracies of 44.23% on the challenge test set through score-level fusion.

TL;DR

Recognizing human emotions in real-world movie clips is notoriously difficult due to environmental noise and occlusion. This paper leverages Extreme Learning Machines (ELM)—a fast learning paradigm—to process audio-visual data. By focusing on the inner facial regions and employing Kernel ELM, the authors achieved a 44.23% accuracy on the AFEW dataset, proving that smart feature selection and fast solvers can compete with more complex models.

Context: Moving from the Lab to the "Wild"

Prior to the EmotiW challenges, most emotion recognition research focused on "clean" data (e.g., actors in studios). However, real life is messy. The Acted Facial Expression in the Wild (AFEW) dataset uses movie clips to simulate this complexity. The researchers identified that standard SVMs and DNNs, while powerful, are often slow to iterate. They turned to ELMs to explore more hypotheses—like which part of the face matters most—using moderate hardware.

Methodology: The Power of Extreme Learning Machines

The core of this work is the Extreme Learning Machine (ELM). Unlike traditional back-propagation networks, ELM randomly generates or uses PCA for the input layer weights and solves the output weights using a simple Least Squares solution.

1. Visual Strategy: Why the "Inner Face"?

The authors hypothesized that the full face contains "cluttered" information—hair, backgrounds, or occlusions at the edges. By dividing the face into 16 regions (4x4) and selecting only the "inner face" (eyes and mouth), they reduced dimensionality while focusing on regions where emotional actions (Action Units) occur.

Facial Region Selection Figure 1: Comparison of different facial region groupings tested for robustness.

2. Acoustic Strategy: Quality over Quantity

The team tried augmenting the data with other emotional corpora (like EMODB). Surprisingly, this didn't help much because those corpora were recorded in "laboratory" conditions, which are too different from movie data. Instead, using Kernel ELMs and CCA-based feature selection proved more effective.

Experimental Insights

The results demonstrated that Kernel ELMs (RBF) consistently outperformed basic ELMs. A crucial discovery was that PCA-initialized input weights provided a substantial boost over random initialization, suggesting that even in "random" networks, a data-driven starting point helps significantly.

Results Table Table 1: Performance comparison between Basic ELM, PCA-ELM, and Kernel versions.

In the final test set, the multimodal fusion (combining audio and video scores) reached 44.23%, a significant jump from the individual modalities.

Critical Analysis & Conclusion

Takeaway

This paper serves as a reminder that feature engineering and fast learning algorithms still have a place in the era of Deep Learning. By focusing on the "inner face," the authors built a system that is naturally more resilient to the "wild" conditions of the dataset.

Limitations

The system struggled with specific classes like Disgust and Surprise in the audio modality (as seen in the confusion matrix). This indicates that the current acoustic features (MFCCs, F0) might not capture the nuances of these specific emotions as well as visual features do.

Future Outlook

While this work used LBP-TOP features (state-of-the-art in 2014), the logic of using ELM as a fast backend for multimodal fusion remains highly relevant today for edge computing and real-time emotion sensing where GPU resources are limited.

Confusion Matrix Figure 2: Final Fusion System Confusion Matrix showing strengths in Anger and Neutral classes.

Find Similar Papers

Try Our Examples

  • Find recent papers from 2020-2024 that utilize Extreme Learning Machines or their variants for multimodal affective computing in uncontrolled environments.
  • Which paper first introduced the Acted Facial Expression In The Wild (AFEW) dataset, and how has the SOTA accuracy evolved in the EmotiW challenges over the last decade?
  • Explore how contemporary Vision Transformer (ViT) architectures incorporate the "inner face" spatial bias similar to the regional selection method proposed in this study.
Contents
ELM for Emotion Recognition: Speed and Accuracy in the Wild
1. TL;DR
2. Context: Moving from the Lab to the "Wild"
3. Methodology: The Power of Extreme Learning Machines
3.1. 1. Visual Strategy: Why the "Inner Face"?
3.2. 2. Acoustic Strategy: Quality over Quantity
4. Experimental Insights
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook