C-GCN: Decoding Human Emotion via Inter-Video Correlations
11137_C-GCN Correlation Based Graph Convolutional Network for Audio-Video Emotion Recognition.
The paper introduces C-GCN, a Correlation-based Graph Convolutional Network for audio-video emotion recognition. It leverages multi-head attention to model inter-class and intra-class relationships between video clips, achieving SOTA results on AFEW (63.55%) and eNTERFACE 05 (97.07%).
TL;DR
Recognizing emotions from audio-visual data is notoriously difficult due to "weak expressions" and "conflicting emotional states." While most researchers focus on fusing audio and video within a single clip, C-GCN (Correlation-based Graph Convolutional Network) shifts the paradigm. By treating video clips as nodes in a global graph and using Multi-Head Attention to learn the edges, the model "consults" other samples to refine its own predictions, achieving state-of-the-art accuracy with remarkable inference speed.
The Missing Piece: Context Beyond the Clip
In traditional Automatic Emotion Recognition (AER), the model is an island. It looks at Video A and tries to decide if the person is "Sad" or "Neutral" based solely on the pixels and audio waves of Video A.
However, emotions are abstract and relative. The authors of C-GCN argue that inter-video correlations—how Video A relates to Video B (even if they are from different people)—are crucial. Current methods suffer because:
- They ignore global information across the dataset.
- Subtle "weak emotions" (like disgust or surprise) are often misclassified because the model lacks a comparative reference within the feature space.
Methodology: Building the Emotion Graph
The C-GCN pipeline follows a sophisticated tripartite architecture:
1. Robust Feature Extraction
The model uses a dual-flow system:
- Audio Flow: Spectrograms are processed via a Fully Convolutional Network (FCN) with an attention mechanism to focus on informative time-frequency units.
- Image Flow: Faces are tracked using dlib, features extracted via FR-Net-B (fine-tuned on FER2013), and temporal dynamics captured via BiLSTM.
- Fusion: The audio and visual vectors are integrated using Factorized Bilinear Pooling (FBP), which captures complex associations more effectively than simple concatenation.
2. Graph Generation with Multi-Head Attention
This is the core innovation. Instead of a static graph, C-GCN uses:
- Initial Edges: Based on category labels (for training) and cosine similarity (for testing).
- Multi-Head Attention: This transforms the sparse initial graph into a set of fully connected edge-weighted graphs. This step eliminates "isolated nodes" and allows the model to predict hidden relationships between different emotion classes.
Fig 1. The overall architecture of C-GCN, illustrating the flow from raw data to graph-based classification.
3. Feature Updating via Dense GCN
To prevent the vanishing gradient problem and the loss of original features, the model employs Densely Connected GCNs. Each layer receives the concatenation of all previous layers' outputs. The information from multiple attention heads is finally fused via mean-pooling to create a highly discriminative video descriptor.
Experimental Results: SOTA Performance
C-GCN was tested on the AFEW and eNTERFACE 05 datasets.
- Efficiency: Despite the graph complexity, C-GCN reached a classification speed of 0.12s per video, outperforming competitors like Zhou et al. and Hu et al.
- Accuracy: On AFEW, it achieved 63.55%, an impressive feat for a single-model approach compared to multi-model ensembles.
- The "Weak Emotion" Breakthrough: The model showed significant improvements in recognizing "Disgust" and "Surprise," which are traditionally the hardest categories to crack.
Table 1. Ablation study showing the impact of FBP and Multi-Head Attention on final accuracy.
Critical Insight: Why Does It Work?
The genius of C-GCN lies in how it handles Manifold Learning. By using Multi-Head Attention to predict edges, the model essentially learns the underlying structure of the "emotion manifold." If Video A is a "weak" version of Happy, the graph creates a strong edge to Video B (a "strong" Happy), allowing the features of Video B to "pull" Video A into the correct classification zone during the GCN message-passing stage.
Conclusion & Future Outlook
C-GCN successfully demonstrates that "no video is an island." By leveraging the hidden correlations between different samples, the model achieves a more nuanced understanding of affect.
Future Work: The authors suggest the next step is modeling cross-modal correlations—using GCNs to map the interaction between audio features and visual landmarks directly, potentially further reducing the confusion between similar emotional states.
Keywords: Emotion Recognition, GCN, Multi-modal Fusion, Multi-head Attention, Deep Learning.
