Beyond the Face: Enhancing Continuous Emotion Recognition via Head Pose and Eye Gaze

Continuous Emotion Recognition in Videos by Fusing Facial Expression, Head Pose and Eye Gaze

2019-10-14
Suowei Wu, Zhengyin Du, Weixin Li, Di Huang, Yunhong Wang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a novel multimodal framework for continuous emotion recognition in videos by fusing facial expressions with head pose and eye gaze (P&G) clues. By employing a Temporal Convolutional Network (TCN) and a P&G-guided attention mechanism, the method achieves state-of-the-art performance on the RECOLA dataset, particularly in predicting valence and arousal.

TL;DR

Researchers from Beihang University have developed a new deep learning framework that proves we shouldn't just look at a person's face to understand their feelings. By fusing facial expressions with Head Pose and Eye Gaze (P&G) clues through a specialized attention mechanism, they achieved a new SOTA on the RECOLA dataset, demonstrating that how we hold our head and where we look provides vital context for interpreting facial cues.

The Problem: The "Frontal Face" Bias

In the world of affective computing, the face is king. Most models are trained on near-frontal views where every micro-expression is visible. However, real-world Human-Computer Interaction (HCI) is messy. People bow their heads, look away in shame, or tilt their heads in anger.

Current SOTA methods face two major hurdles:

  1. Noise Sensitivity: Changes in head pose bring significant noise to facial feature extraction. If the model can't see the face clearly, the "emotion" it detects is likely hallucinated.
  2. Information Neglect: Head pose and eye gaze are emotional signals in their own right (e.g., a bowed head signalizing sadness or submissiveness), but they are rarely exploited as independent channels of information.

Methodology: Guiding and Augmenting

The authors propose a dual-pathway approach to solve these issues, integrated into an end-to-end architecture featuring ResNet18 with Squeeze-and-Excitation (SE) blocks for spatial features and Temporal Convolutional Networks (TCN) for time-series regression.

1. P&G Guided Temporal Attention

Instead of treating all video frames as equally important, the model uses P&G data to calculate an "importance score." If a subject bows their head, the attention mechanism recognizes that the facial features from those frames are less "credible" and reduces their weight in the final emotional prediction.

Overall Architecture

2. Feature Augmentation

The researchers found a direct correlation between the "pitch" of a head and the "Valence" (positivity/negativity) of an emotion. Therefore, they added an auxiliary line that extracts high-level representations of P&G to be concatenated with facial features, effectively augmenting the model's emotional vocabulary.

Evidence of Success: SOTA Results

The framework was tested on the RECOLA dataset, using the Concordance Correlation Coefficient (CCC) as the primary metric.

MethodArousal (CCC)Valence (CCC)
Khorrami et al. (CNN+LSTM)0.5440.506
Lee et al. (3D-CNN+LSTM)-0.546
Ours (FER-P&G-Net)0.6030.686

The ablation study revealed a fascinating insight: While the attention mechanism significantly boosted Valence (improving it from 0.609 to 0.674), the feature augmentation was more effective for Arousal (improving it from 0.524 to 0.565). Combining both provided the ultimate performance leap.

Attention Visualization

In the visualization above, the model assigns different weights (colors) to frames based on head movements (the black curve). When the subject bows, the model shifts focus to more reliable previous frames.

Critical Insight & Conclusion

The true value of this work lies in its Heuristic Intuition. By recognizing that P&G data serves a dual purpose—both as a "reliability gate" for the face and as an "independent signal"—the authors have created a blueprint for more robust multimodal AI.

Future Outlook: While highly effective, the model currently relies on OpenFace 2.0 for ROI extraction. Future iterations could benefit from integrating raw gaze heatmaps or 3D head meshes directly into the end-to-end pipeline to capture even finer nuances of "non-verbal" emotional leakage.

Find Similar Papers

Try Our Examples

  • Search for recent papers on multimodal emotion recognition that combine facial expressions with physiological signals like ECG or EDA on the RECOLA dataset.
  • Which study first introduced the use of Temporal Convolutional Networks (TCN) for regression in affective computing, and how does this paper's implementation differ?
  • Explore how head pose and eye gaze features are being used in zero-shot or cross-dataset emotion recognition to handle domain shift in facial appearances.
Contents
Beyond the Face: Enhancing Continuous Emotion Recognition via Head Pose and Eye Gaze
1. TL;DR
2. The Problem: The "Frontal Face" Bias
3. Methodology: Guiding and Augmenting
3.1. 1. P&G Guided Temporal Attention
3.2. 2. Feature Augmentation
4. Evidence of Success: SOTA Results
5. Critical Insight & Conclusion