KAVAN: Decoding Emotions in GIFs through Human-Centered Attention
Human-Centered Emotion Recognition in Animated GIFs
This paper introduces the Keypoint Attended Visual Attention Network (KAVAN) for human-centered emotion recognition in animated GIFs. By leveraging a facial attention module and a Hierarchical Segment LSTM (HS-LSTM), the method achieves SOTA performance on the MIT GIFGIF dataset through multi-task learning for both classification and intensity regression.
TL;DR
Animated GIFs have become the lingua franca of digital emotion, yet AI usually treats them as mere "short videos." Researchers from the University of Rochester have introduced KAVAN (Keypoint Attended Visual Attention Network), a framework that prioritizes human facial expressions and hierarchical temporal structures. By using noisy facial keypoints as a guide rather than a crutch, KAVAN achieves new SOTA results on the MIT GIFGIF dataset, proving that in the world of GIFs, the face is the window to the soul.
Problem & Motivation: Why GIFs are Not Just "Small Videos"
GIFs possess two unique properties that traditional Computer Vision models often ignore:
- Human-Centricity: Over 50% of GIFs feature human faces, and most others feature personified characters. Traditional CNNs often get distracted by background clutter.
- Temporal Density: Unlike long videos, GIFs have no "filler" frames. Every frame is highly salient. Standard LSTMs tend to lose information from early frames by the time they reach the end of the sequence.
Previous attempts to use facial keypoints were brittle—if the keypoint detector failed (common in low-res or stylized GIFs), the whole model failed. The authors set out to create a system that is robust to "missing" data and captures the essence of temporal evolution.
Methodology: The KAVAN Architecture
KAVAN's innovation lies in two distinct modules: the Facial Soft Attention Module and the HS-LSTM.
1. Robust Facial Attention
Instead of feeding keypoint coordinates directly into the network, KAVAN treats them as supervision. The model learns to predict a facial region mask () by comparing its internal attention to a heatmap generated from estimated keypoints.
- The "Why": If a keypoint detector has low confidence (common in blurry GIFs), the supervision's weight is lowered. This makes the model "smart" enough to find faces even when the keypoint labels are wrong or missing, as seen in cartoon characters.
Fig 1: Overall structure of KAVAN featuring the attention module (blue) and the temporal module.
2. HS-LSTM: Temporal Heirarchy
To solve the "forgetting" problem, the Hierarchical Segment LSTM (HS-LSTM) splits a GIF into segments.
- Tier 1: Learns coarse representations for each segment.
- Tier 2: Uses its own frames plus the knowledge from Tier 1 to produce a refined global representation. This ensures that even the very first frame of a 2-second GIF contributes significantly to the final classification.
Fig 2: A two-tier HS-LSTM learning features from coarse-to-fine resolution.
Experiments & Results: Performance That Matters
The researchers tested KAVAN on the MIT GIFGIF dataset, targeting 17 distinct emotions (e.g., contempt, relief, pride).
Quantitative SOTA
The results show a clear additive benefit of each module:
- Baseline (ResNet-50 + LSTM): 61.47% Accuracy
- With Soft Attention: +2.08%
- With HS-LSTM: +1.08%
- Multi-Task Learning (MTL): By training for both category classification and emotion intensity regression simultaneously, the model reached 68.27% accuracy.
Qualitative Interpretability
Perhaps most impressively, KAVAN demonstrates "cross-domain" intelligence. In Fig 3 below, we see that even though the keypoint detector might provide poor data for a cartoon character, the Attention Mask (lower row) successfully isolates the character's face, proving the model has learned the semantic concept of a face.
Fig 3: Visualization of the attention masks. Note how the model focuses on the facial region even in artistic/cartoon styles.
Critical Insight & Conclusion
KAVAN proves that soft supervision is often superior to hard feature fusion. By teaching the model where to look (using keypoints as a hint) rather than what the keypoints are, the authors created a robust system that handles the chaotic variety of social media content.
Future Outlook: While KAVAN is localized to GIFs, its hierarchical temporal approach could be revolutionary for other "dense" video tasks, such as sign language recognition or surgical video analysis, where every millisecond counts.
