Beyond Regions: Improving Image Emotion Recognition via Weakly Supervised Intensity Learning
Weakly Supervised Emotion Intensity Prediction for Recognition of Emotions in Images
This paper introduces a novel end-to-end deep neural network for image emotion recognition that leverages weakly supervised emotion intensity learning. By integrating an Intensity Prediction Stream built on a Feature Pyramid Network (FPN) with two classification streams, the method achieves new SOTA results on benchmarks like FI-8 and Emotion-6.
TL;DR
Recognizing emotions in images is inherently difficult because emotions are subjective and rarely occupy the entire frame. This paper introduces an end-to-end framework that doesn't just classify an image, but predicts an Emotion Intensity Map to pinpoint exactly where the "feeling" comes from. By using a weakly supervised approach that transforms class activation into intensity guidance, the authors achieved SOTA performance on the FI-8 and Emotion-6 datasets.
The "Weak Label" Bottleneck
In standard image classification (like ImageNet), a "cat" is usually a concrete object. In emotion recognition, a "sad" label might only apply to a small drooping flower in a vast neutral landscape. Most existing datasets only provide image-level labels, which are "weak" because they don't tell the model where to look.
Prior attempts to solve this used:
- Handcrafted features: Limited by the expert's imagination.
- Region proposals: Computationally heavy and often required manual bounding boxes for pre-training.
The authors' insight was simple yet powerful: Emotion is a continuous intensity field, not just a binary box.
Methodology: The Triple-Stream Architecture
The proposed network consists of three distinct yet cooperative streams:
- The Probing Stream (First Classification): A standard CNN that generates initial class activation maps (CAM). These maps serve as "pseudo-ground truth" for what regions drive the emotion.
- The Discovery Stream (Intensity Prediction): Built on a Feature Pyramid Network (FPN), this stream learns to predict the intensity map directly from the image. It uses a combination of three losses:
- RMSEL: For pixel-wise intensity accuracy.
- Gradient Loss: To keep the edges of emotional regions sharp.
- Surface Normal Loss: To ensure the geometric "shape" of the intensity reflects the image structure.
- The Refinement Stream (Second Classification): This stream takes the original features and "multiplies" them by the predicted intensity map. This forces the model to ignore background noise and focus its representation on high-intensity emotional stimuli.
Figure: The three-stream architecture showing the flow from Pseudo-Intensity generation to final classification.
Why the FPN Matters
By using an FPN (shown below), the network can extract multilevel features. This is critical because some emotional cues are small (a subtle facial expression), while others are global (the color palette of a sunset). The FPN's top-down pathway combines these semantic scales into a single, high-resolution intensity map.
Figure: The Intensity Prediction Subnetwork built on Top of FPN.
Experimental Results & Insights
The results across datasets were consistently superior to vanilla architectures and previous region-based SOTA like Rao et al.
- FI-8 Dataset: Jumped from 66.16% (Vanilla ResNet-101) to 75.91%.
- Emotion-6: Achieved 60.41%, noticeably better than much larger models like ResNet-152.
One fascinating takeaway from the ablation studies is that the "Surface Normal Loss" and "Gradient Loss" (typically used in depth estimation) significantly helped. This suggests that the spatial "contours" of an emotion are just as important as the raw pixel values.
Figure: Visualization comparing CAM-generated maps (left) vs. the model's Predicted Intensity Maps (right).
Critical Analysis & Future Outlook
While the method is robust, the confusion matrices reveal that "Fear" and "Anger" remain difficult to distinguish—often because these emotions share similar visual triggers in the wild.
Future Directions:
- Cross-Modal Guidance: Could textual metadata (tags/comments) be used to further refine the pseudo-intensity maps?
- Temporal Stability: Applying this intensity-based logic to video clips, where emotion intensity fluctuates over time.
In conclusion, this work proves that we don't need expensive manual annotations to understand "where" an emotion is. By teaching a network to predict its own attention maps, we get a model that is both more accurate and more interpretable.
