Contemplating Visual Emotions: Overcoming Dataset Bias via Curriculum Web-Learning

Contemplating Visual Emotions: Understanding and Overcoming Dataset Bias

2018-01-01
Rameswar Panda, Jianming Zhang, Haoxiang Li, Joon-Young Lee, Xin Lu, Amit K. Roy-Chowdhury
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces WEBEmo, a massive-scale dataset for visual emotion recognition, and proposes a Curriculum Guided Training strategy to mitigate significant dataset biases. By leveraging 268,000 weakly labeled stock images across a 25-category hierarchy, the authors achieve SOTA performance on multiple image and video emotion benchmarks without manual labeling.

Executive Summary

TL;DR: Most AI models for emotion recognition don't actually "feel" the emotion; they just recognize objects commonly associated with them in biased datasets. This paper exposes this "Dataset Bias" and introduces WEBEmo, a dataset of 268,000 images, alongside a Curriculum Guided Training strategy that trains models on an evolutionary path from simple to complex emotions.

Positioning: This work is a critical "reality check" for the affective computing community. It shifts the focus from chasing benchmark scores to fundamental model generalization and "unbiased" feature learning.


The "Sad Theme Park" Problem: Deep-Diving into Bias

Standard emotion datasets often suffer from what the authors call Positive and Negative Set Bias.

For example, if a dataset only contains images of "Amusement Parks" under the category "Joy," a CNN will learn that "Rollercoaster = Joy." When shown a photo of an abandoned, derelict theme park (which evokes "Sadness"), the model provides a 99.9% confidence score for "Joy" simply because it sees the tracks. This is not emotion recognition; it's object correlation.

The authors quantified this using a Name That Dataset Game, proving that a simple classifier can tell which dataset an image belongs to with 63.67% accuracy—meaning datasets have "signatures" or "biases" that have nothing to do with the actual labels.


Methodology: High-Volume Data & Curriculum Learning

To solve this, the authors moved away from small, manually labeled sets to a Webly Supervised approach.

1. The WEBEmo Dataset

By crawling stock photo sites using keywords derived from Parrott’s Emotion Hierarchy, they amassed 268,000 high-quality images. This diversity ensures the model sees a "Sad Balloon" or a "Joyful Rain," breaking the rigid object-emotion correlations found in smaller sets.

2. Curriculum Guided Training

Learning fine-grained emotions (e.g., exasperation vs. disappointment) is hard. The authors leveraged a psychological hierarchy to guide the CNN's learning process:

  • Stage 1 (Coarse): Binary classification (Positive vs. Negative).
  • Stage 2 (Mid): 6 Secondary emotions (Anger, Fear, Joy, Love, Sadness, Surprise).
  • Stage 3 (Fine): 25 Tertiary fine-grained emotions.

This "easy-to-hard" transition acts as a biological-like regularizer, helping the model learn stable features before tackling the noise of fine-grained labels.

Model Architecture and Hierarchy Figure: The 3-level emotion hierarchy used for Curriculum Learning.


Experiments and Results

The most striking result is the Binary Cross-Dataset Generalization. When a model trained on WEBEmo was tested on foreign datasets (like Emotion-6), it actually outperformed models trained on those specific datasets (81.41% vs 77.72%). This proves that "Diversity > Manual Labeling."

Quantitative Impact:

  • Accuracy Boost: Improved SOTA on the Deep Sentiment dataset by 8%.
  • Negative Bias: As shown in the table below, the WEBEmo-trained model's performance dropped by nearly 0% when introduced to external negative samples, whereas previous models dropped by ~25%.
Train DatasetTest on Others (Mean Acc)Drop in Performance
Deep Emotion63.52%High
WEBEmo (Ours)72.76%Minimal

Correlation Analysis Figure: Correlation analysis showing WEBEmo (right) has much higher conditional entropy (less bias) compared to Deep Emotion (left).


Critical Insight & Conclusion

The true value of this paper isn't just a new SOTA; it’s the demonstration that Emotion is a feature, not just a label.

By applying these "Emotion Features" to Video Summarization, they achieved a 3% mAP increase. This suggests that AI systems for summarization, advertising, and content creation should prioritize emotional cues over raw pixel geometry.

Limitations: While stock images are high-quality, they are "composed." Future work could look into "candid" or "raw" social media data to capture even more visceral emotional expressions.

Summary Takeaway: Start with the hierarchy. If your model can't tell "Happy" from "Sad" across different datasets, it has no business trying to label "Melancholy."

Find Similar Papers

Try Our Examples

  • Search for recent papers that address "dataset bias" specifically in the context of affective computing or subjective visual tasks beyond object recognition.
  • Which original psychology paper first defined the "Parrott's wheel of emotions," and how have other deep learning architectures (like Transformers) incorporated this hierarchy recently?
  • Find studies that integrate emotional features or WEBEmo-based pretrained weights into multi-modal tasks such as text-to-image generation or affective video captioning.
Contents
Contemplating Visual Emotions: Overcoming Dataset Bias via Curriculum Web-Learning
1. Executive Summary
2. The "Sad Theme Park" Problem: Deep-Diving into Bias
3. Methodology: High-Volume Data & Curriculum Learning
3.1. 1. The WEBEmo Dataset
3.2. 2. Curriculum Guided Training
4. Experiments and Results
4.1. Quantitative Impact:
5. Critical Insight & Conclusion