Beyond the Face: Enhancing Continuous Emotion Recognition via Head Pose and Eye Gaze
Continuous Emotion Recognition in Videos by Fusing Facial Expression, Head Pose and Eye Gaze
This paper introduces a novel multimodal framework for continuous emotion recognition in videos by fusing facial expressions with head pose and eye gaze (P&G) clues. By employing a Temporal Convolutional Network (TCN) and a P&G-guided attention mechanism, the method achieves state-of-the-art performance on the RECOLA dataset, particularly in predicting valence and arousal.
TL;DR
Researchers from Beihang University have developed a new deep learning framework that proves we shouldn't just look at a person's face to understand their feelings. By fusing facial expressions with Head Pose and Eye Gaze (P&G) clues through a specialized attention mechanism, they achieved a new SOTA on the RECOLA dataset, demonstrating that how we hold our head and where we look provides vital context for interpreting facial cues.
The Problem: The "Frontal Face" Bias
In the world of affective computing, the face is king. Most models are trained on near-frontal views where every micro-expression is visible. However, real-world Human-Computer Interaction (HCI) is messy. People bow their heads, look away in shame, or tilt their heads in anger.
Current SOTA methods face two major hurdles:
- Noise Sensitivity: Changes in head pose bring significant noise to facial feature extraction. If the model can't see the face clearly, the "emotion" it detects is likely hallucinated.
- Information Neglect: Head pose and eye gaze are emotional signals in their own right (e.g., a bowed head signalizing sadness or submissiveness), but they are rarely exploited as independent channels of information.
Methodology: Guiding and Augmenting
The authors propose a dual-pathway approach to solve these issues, integrated into an end-to-end architecture featuring ResNet18 with Squeeze-and-Excitation (SE) blocks for spatial features and Temporal Convolutional Networks (TCN) for time-series regression.
1. P&G Guided Temporal Attention
Instead of treating all video frames as equally important, the model uses P&G data to calculate an "importance score." If a subject bows their head, the attention mechanism recognizes that the facial features from those frames are less "credible" and reduces their weight in the final emotional prediction.

2. Feature Augmentation
The researchers found a direct correlation between the "pitch" of a head and the "Valence" (positivity/negativity) of an emotion. Therefore, they added an auxiliary line that extracts high-level representations of P&G to be concatenated with facial features, effectively augmenting the model's emotional vocabulary.
Evidence of Success: SOTA Results
The framework was tested on the RECOLA dataset, using the Concordance Correlation Coefficient (CCC) as the primary metric.
| Method | Arousal (CCC) | Valence (CCC) |
|---|---|---|
| Khorrami et al. (CNN+LSTM) | 0.544 | 0.506 |
| Lee et al. (3D-CNN+LSTM) | - | 0.546 |
| Ours (FER-P&G-Net) | 0.603 | 0.686 |
The ablation study revealed a fascinating insight: While the attention mechanism significantly boosted Valence (improving it from 0.609 to 0.674), the feature augmentation was more effective for Arousal (improving it from 0.524 to 0.565). Combining both provided the ultimate performance leap.

In the visualization above, the model assigns different weights (colors) to frames based on head movements (the black curve). When the subject bows, the model shifts focus to more reliable previous frames.
Critical Insight & Conclusion
The true value of this work lies in its Heuristic Intuition. By recognizing that P&G data serves a dual purpose—both as a "reliability gate" for the face and as an "independent signal"—the authors have created a blueprint for more robust multimodal AI.
Future Outlook: While highly effective, the model currently relies on OpenFace 2.0 for ROI extraction. Future iterations could benefit from integrating raw gaze heatmaps or 3D head meshes directly into the end-to-end pipeline to capture even finer nuances of "non-verbal" emotional leakage.
