FullExpression: Bringing Real-Time Emotion Recognition to the Browser via Transfer Learning

FullExpression Using Transfer Learning in the Classification of Human Emotions

2020-09-10
Ricardo Rocha, Isabel Praça
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents "FullExpression," a facial emotion classification system utilizing Transfer Learning on a Deep Learning model. By fine-tuning a MobileNet V1 architecture, the authors achieved an 80% accuracy in classifying the seven universal emotions within a browser-based environment.

TL;DR

FullExpression is a research-driven ecosystem that bridges the gap between complex Deep Learning (DL) models and accessible web applications. By utilizing Transfer Learning on a MobileNet V1 backbone, the authors developed a system capable of identifying seven core human emotions with ~80% accuracy directly within a web browser using TensorFlow.js.

Problem & Motivation

Human-Computer Interaction (HCI) is becoming increasingly emotional. However, the industry faces a dilemma: traditional FACS (Facial Action Coding System) methods are often rigid and sensitive to lighting or ethnicity, while modern DL models are typically too "heavy" to run without specialized server-side hardware.

The authors observed that most state-of-the-art (SOTA) works achieve high accuracy (up to 98% on controlled datasets like CK+) but fail to provide a practical, cross-platform tool. Their insight was to leverage MobileNet, an architecture designed specifically for resource-constrained environments, and adapt it for facial expressions using a hybrid dataset.

Methodology: The Core

The "FullExpression" ecosystem is built on a modular architecture that separates concern between UI, data handling, and the DL model.

1. The Model Architecture

Instead of building a model from scratch, the team used Transfer Learning. They took a pre-trained MobileNet V1 and performed the following modifications:

  • Input Layer: Set to 224 x 224 x 3.
  • Bottleneck: Retained the 89-layer structure to leverage features already learned from large-scale image sets.
  • Output Layer: Replaced the final classification layer with a 1 x 1 x 7 layer, representing the seven universal emotions (Joy, Sadness, Surprise, Fear, Anger, Disgust, and Contempt).

2. The Browser Pipeline

To achieve real-time inference, the application follows a strict pipeline:

  1. Detection: Face extraction via the Viola-Jones algorithm.
  2. Normalization: Resizing to 300x300 and grayscale conversion to reduce dimensionality.
  3. Inference: Running the model via TensorFlow.js, utilizing the client's local GPU/CPU.

FullExpression Architecture

Experiments & Results

The model was trained on a combined database of 4,699 images (including KDEF, TFEID, and JAFFE).

SOTA Comparison

FullExpression achieved a performance of ~78.4%, which is highly competitive considering it is a browser-based solution. Notably, it outperformed EmotionDan (~75% accuracy) despite using significantly less training data, proving that data quality and transfer learning efficiency can outweigh sheer volume.

Performance by Emotion

The study revealed a common "bias" in emotion AI:

  • "Happy" was identified with near-perfect accuracy (97.8%).
  • "Afraid" proved the most elusive (62.5%), often confused with surprise or sadness.

Accuracy Results across Databases

Critical Analysis & Conclusion

The core achievement of this paper isn't necessarily a new "highest score" in accuracy, but a proof-of-concept for deployment. By moving the inference to the browser, the authors address privacy concerns (images don't need to be uploaded to a server) and reduce infrastructure costs.

Takeaways

  1. Efficiency over Size: A well-tuned lightweight model (MobileNet) can outperform larger models if the transfer learning process is handled correctly.
  2. Transfer Learning is Key: Retraining all 89 layers of MobileNet for a niche task like emotion recognition allows the model to adapt generic visual features to subtle muscle movements.

Limitations & Future Work

The accuracy drop on the "Face Place" database (down to 43%) suggests that the model still struggles with heavily disguised or highly varied facial structures. Future research should likely move toward Multi-modal Affective Computing, combining facial analysis with voice or physiological signals (heart rate, skin conductance) to clear up the ambiguity between overlapping emotions like "Anger" and "Disgust."

Find Similar Papers

Try Our Examples

  • Find recent papers that optimize MobileNet or other lightweight CNNs for real-time facial expression recognition on edge devices or browsers.
  • Which paper first proposed the MobileNet V1 architecture, and how has its depthwise separable convolution influenced subsequent transfer learning techniques in computer vision?
  • Explore research that integrates physiological responses or body language with facial expression data to improve the multi-modal classification of human emotions.
Contents
FullExpression: Bringing Real-Time Emotion Recognition to the Browser via Transfer Learning
1. TL;DR
2. Problem & Motivation
3. Methodology: The Core
3.1. 1. The Model Architecture
3.2. 2. The Browser Pipeline
4. Experiments & Results
4.1. SOTA Comparison
4.2. Performance by Emotion
5. Critical Analysis & Conclusion
5.1. Takeaways
5.2. Limitations & Future Work