Personalized FER: Bridging the Gap Between Universal Models and Individual Expression

Personalized models for facial emotion recognition through transfer learning

2020-08-13
Martina Rescigno, Matteo Spezialetti, Silvia Rossi
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a transfer learning framework for personalized Facial Emotion Recognition (FER) using a dimensional model (Valence-Arousal). By fine-tuning an AlexNet model pre-trained on the large-scale "in-the-wild" AffectNet dataset with small amounts of subject-specific data from the AMIGOS dataset, the authors achieve high recognition performance (RMSE of 0.09 for Valence and 0.1 for Arousal).

TL;DR

Building a facial emotion recognition (FER) system that works for everyone is hard because we all express feelings differently. This paper proposes a transfer learning strategy: take a deep brain (AlexNet) already trained on millions of internet faces (AffectNet) and give it a "personal touch" by fine-tuning it on a few hundred photos of a specific person. The result? A massive jump in accuracy (RMSE ~0.1) with very little data required from the user.

The "One-Fits-All" Fallacy

In the world of Affective Computing, most models try to be universal. However, factors like bone structure, cultural background, and even lighting mean that a "happy" face for one person might look like "neutral" to a model trained on someone else.

The authors identify two failing extremes:

  1. Generalized Models: They miss the nuances of the individual.
  2. Purely Personal Models: They require thousands of labeled images per person to learn from scratch—an impossible task for real-world apps.

The Insight: Use the "In-the-Wild" knowledge (general features like eyes, mouth shapes) and adapt only the final decision-making layers to the specific person.

Methodology: The Transductive Shortcut

The researchers chose AlexNet as their workhorse. The process involves a two-stage pipeline:

  1. Pre-training: Training on AffectNet, which contains >1 million images. This teaches the model what a human face looks like and the general geometry of emotions.
  2. Personalized Fine-tuning: They "freeze" the convolutional layers (the eyes of the model) and only retrain the dense layers (the brain of the model) using a small slice of data from the AMIGOS dataset.

Model Architecture and Fine-tuning Workflow

Active vs. Passive Sampling

Do we need to pick specific "hard" images to label? The authors tested Greedy Sampling and Monte-Carlo Dropout (MCDUE) to find the most "uncertain" images. Surprisingly, they found that in FER, the transfer of knowledge is so powerful that even random sampling works nearly as well as complex active learning strategies.

Experimental Battleground: Valence vs. Arousal

The study reveals a fascinating psychological split:

  • Valence (Positive vs. Negative): Generalizes well. The model can easily tell a "happy" heart from a "sad" one across different people.
  • Arousal (Calm vs. Excited): Highly personal. How "intense" a face looks depends heavily on the individual's baseline. This is where the personalized data provided the biggest boost.

Performance Comparison: No Transfer vs. Full Transfer

As shown in the charts, the "Transfer 100% Labeling" (green bars) achieved the lowest RMSE, proving that the hybrid approach is superior to using either only general data or only personal data.

Practical Gains: How Much Data is Enough?

One of the most valuable takeaways for developers is the elbow point. The authors found that after labeling about 240 to 320 frames (roughly 20-30% of their small set), the error rate hits a plateau.

Learning Curves and Sample Efficiency

Critical Insight & Future Outlook

This work proves that we don't need "Big Data" for every single user; we need "Smart Transfer."

Key Takeaways:

  • Valence is universal; Arousal is personal.
  • AlexNet, despite being an older architecture, remains highly effective for localized feature extraction in FER.
  • The need for active learning is lower than expected because the pre-trained features are already high-quality.

Future Work: The next step is moving beyond static images to temporal models (LSTMs/Transformers) to see how personalization handles sequences of expressions over time.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Vision Transformers (ViT) instead of CNNs for personalized facial emotion recognition via transfer learning.
  • Which paper first proposed the Valence-Arousal circumplex model of affect, and how has its application in deep learning evolved since Russell (1980)?
  • Explore studies investigating the application of Meta-Learning (e.g., MAML) for fast adaptation in few-shot facial expression recognition tasks.
Contents
Personalized FER: Bridging the Gap Between Universal Models and Individual Expression
1. TL;DR
2. The "One-Fits-All" Fallacy
3. Methodology: The Transductive Shortcut
3.1. Active vs. Passive Sampling
4. Experimental Battleground: Valence vs. Arousal
5. Practical Gains: How Much Data is Enough?
6. Critical Insight & Future Outlook