Subtraction Pre-Processing: Redefining Temporal Dynamics in Emotion Recognition

Human Emotion Recognition in Video Using Subtraction Pre-Processing

2019-02-22
Zhihao He, Tian Jin, Amlan Basu, John J. Soraghan, Gaetano Di Caterina, Lykourgos Petropoulakis
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a frame-subtraction pre-processing method combined with Convolutional Neural Networks (CNN) for video-based human emotion recognition. By focusing on dynamic pixel changes between frames, the system achieved a state-of-the-art accuracy of 79.74% on the RAVDESS dataset using the AlexNet architecture.

TL;DR

Recognizing human emotion in video typically requires heavy-duty sequence modeling. This paper challenges that norm by introducing a Subtraction Pre-Processing method. By calculating the difference between aligned frames, the researchers isolate the "motion features" of an emotion, allowing a standard AlexNet to outperform human raters with a 79.74% accuracy on the RAVDESS dataset.

Motivation: Why Background is Noise

In the context of facial expression, the static parts of a video—the background, the neck, and even the hair—are often irrelevant to the actual emotion being conveyed. Traditional 2D-CNNs struggle because they treat every pixel with equal importance, and 3D-CNNs/LSTMs are computationally expensive.

The authors' core intuition is that emotion is change. By subtracting frame from frame , static pixels become zero (black), and only the relative movement of facial muscles remains. This creates a "feature-rich" sparse image that highlights exactly what the neural network needs to see.

Methodology: The Art of the Gap

The system follows a rigorous pipeline:

  1. Face Detection & Alignment: Using Haar-like features to find the face and rotating it based on eye/mouth coordinates to ensure precise pixel-to-pixel correspondence.
  2. Subtraction Operation: This isn't just neighboring frame subtraction. The authors introduce a "Gap" (distance between frames) and "Stride" (sampling rate).
    • If the gap is too small, the difference is negligible.
    • If the gap is too large, the temporal relationship is lost.
  3. CNN Classification: The resulting "difference images" are fed into architectures like AlexNet, GoogleNet, and ResNet.

System Architecture Figure 1: The proposed system flow from video decomposition to CNN classification.

Experiments & Surprising Results

The researchers tested their method on the RAVDESS dataset (vocal expressions and songs). One of the most striking findings was that AlexNet, a relatively shallow model by modern standards, performed the best.

  • AlexNet: 79.74% Accuracy
  • ResNet-4: 75.89% Accuracy
  • GoogleNet: 62.89% Accuracy

The authors hypothesize that because the pre-processing already "simplifies" the data by removing noise, hyper-deep networks like GoogleNet (100+ layers) may actually over-abstract the features, leading to poor performance.

Accuracy Comparison Table 1: Comparison of the proposed model against human raters and other visual/acoustic baselines.

Critical Insight: More is Not Always Better

The study reveals a vital lesson in machine learning: Smart pre-processing can reduce the need for model complexity. By transforming the video into a "difference vision," the researchers effectively performed manual feature engineering that "taught" the AI to focus on muscle contraction rather than identity or lighting.

Limitations

  • Pose Sensitivity: The current method relies heavily on frontal face alignment. If a subject turns their head significantly, the subtraction becomes chaotic.
  • Real-time Constraints: While the CNN inference is fast, the alignment and subtraction steps currently limit smooth live-video analysis without further optimization.

Conclusion

This work demonstrates that "seeing the change" is more efficient than "seeing the whole." By achieving a 4.78% improvement over human accuracy using only visual data, the subtraction pre-processing method provides a lightweight yet powerful alternative to complex recurrent architectures in the field of affective computing.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize motion history images (MHI) or differential frames in deep learning for micro-expression recognition.
  • What are the current SOTA benchmarks for the RAVDESS dataset using multimodal (audio-visual) fusion versus purely visual methods?
  • Explore how frame subtraction techniques are being integrated into event-based camera processing for low-latency gesture or emotion recognition.
Contents
Subtraction Pre-Processing: Redefining Temporal Dynamics in Emotion Recognition
1. TL;DR
2. Motivation: Why Background is Noise
3. Methodology: The Art of the Gap
4. Experiments & Surprising Results
5. Critical Insight: More is Not Always Better
5.1. Limitations
6. Conclusion