MPATN: Bridging the Affective Gap in Zero-shot Video Emotion Recognition
Zero-shot Video Emotion Recognition via Multimodal Protagonist-aware Transformer Network
This paper introduces the Multimodal Protagonist-aware Transformer Network (MPATN) for the novel task of zero-shot video emotion recognition. By identifying "protagonists" through a dynamic emotional attention mechanism and aligning visual/acoustic features via Noise Contrastive Estimation (NCE), the model achieves state-of-the-art performance on YouTube-8/24, VideoStory-P14, and LIRIS-ACCEDE-11 datasets.
TL;DR
Recognizing emotions in video is difficult because human feelings are subjective and fine-grained. This paper presents MPATN, a transformer-based framework that doesn't just look at a video as a sequence of frames, but identifies the "Protagonist"—the person, animal, or object driving the emotional narrative. By combining this visual focus with advanced acoustic features and contrastive learning, the model can recognize emotions it has never seen during training.
Background & Motivation
Standard emotion recognition models are "closed-world"; they only know what they've seen (e.g., happiness, sadness). But psychological theories like Ortony's define up to 22+ complex emotions (hope, shame, gratitude). Gathering labeled videos for every nuance is impossible.
The authors identify two fatal flaws in prior work:
- Context Blindness: Most models treat all pixels or objects equally, ignoring that an emotion is usually triggered by a specific "protagonist."
- The Affective Gap: There is a massive disconnect between "low-level" pixels/audio and "high-level" human sentiment.
Methodology: Who is the Protagonist?
The core innovation of MPATN lies in its Protagonist-aware design. Instead of feeding the whole frame into a black box, the model uses a two-step process:
1. Dynamic Emotional Attention (DEA)
The model detects objects and ranks them based on:
- Relatedness: How much the object interacts with the overall video scene.
- Affectiveness: How "emotionally charged" the object is, determined by querying external sentiment lexicons (Valence-Arousal scores).
2. Protagonist-aware Transformer
Selected protagonists become the Queries in a Transformer. They "look" at the rest of the video (Keys/Values) to understand the temporal context of the emotion.

3. Multimodal Contrastive Learning
Recognizing "Fear" vs "Terror" is hard visually, but audio levels often provide the missing clue. MPATN uses wav2vec2.0 for audio and employs Noise Contrastive Estimation (NCE) to ensure that video and audio embeddings from the same clip are pulled together in a shared "Affective Space."
Experimental Performance
The researchers tested MPATN on four complex datasets, including YouTube-24 (fine-grained categories) and LIRIS-ACCEDE.
Key Results:
- Accuracy Boost: MPATN consistently beat baselines like E2ET and MGAV. In the YouTube-8 6:2 split, it reached 60.50% accuracy, a significant margin over competitors.
- Ablation Success: Removing the "Dynamic Emotional Attention" (Protagonist identification) caused performance to drop, proving that who we watch matters as much as what we see.

Deep Insights & Takeaway
The success of MPATN suggests that Affective Computing is moving toward Semantic Understanding. By leveraging external knowledge bases (the sentiment lexicons) and focusing on key entities (protagonists), we can build models that generalize to the messy, diverse spectrum of human emotion.
Limitations: The model still struggles with LIRIS-ACCEDE-11, where categories are extremely subtle (e.g., "relaxed" vs "content"). Future work will likely need larger pre-trained Vision-Language models to refine these boundaries.
Conclusion
MPATN is a powerful step toward AI that "understands" the story within a video. By focusing on the protagonist, it bridges the gap between raw data and human-like emotional intuition.
