Subtraction Pre-Processing: Redefining Temporal Dynamics in Emotion Recognition
Human Emotion Recognition in Video Using Subtraction Pre-Processing
The paper introduces a frame-subtraction pre-processing method combined with Convolutional Neural Networks (CNN) for video-based human emotion recognition. By focusing on dynamic pixel changes between frames, the system achieved a state-of-the-art accuracy of 79.74% on the RAVDESS dataset using the AlexNet architecture.
TL;DR
Recognizing human emotion in video typically requires heavy-duty sequence modeling. This paper challenges that norm by introducing a Subtraction Pre-Processing method. By calculating the difference between aligned frames, the researchers isolate the "motion features" of an emotion, allowing a standard AlexNet to outperform human raters with a 79.74% accuracy on the RAVDESS dataset.
Motivation: Why Background is Noise
In the context of facial expression, the static parts of a video—the background, the neck, and even the hair—are often irrelevant to the actual emotion being conveyed. Traditional 2D-CNNs struggle because they treat every pixel with equal importance, and 3D-CNNs/LSTMs are computationally expensive.
The authors' core intuition is that emotion is change. By subtracting frame from frame , static pixels become zero (black), and only the relative movement of facial muscles remains. This creates a "feature-rich" sparse image that highlights exactly what the neural network needs to see.
Methodology: The Art of the Gap
The system follows a rigorous pipeline:
- Face Detection & Alignment: Using Haar-like features to find the face and rotating it based on eye/mouth coordinates to ensure precise pixel-to-pixel correspondence.
- Subtraction Operation: This isn't just neighboring frame subtraction. The authors introduce a "Gap" (distance between frames) and "Stride" (sampling rate).
- If the gap is too small, the difference is negligible.
- If the gap is too large, the temporal relationship is lost.
- CNN Classification: The resulting "difference images" are fed into architectures like AlexNet, GoogleNet, and ResNet.
Figure 1: The proposed system flow from video decomposition to CNN classification.
Experiments & Surprising Results
The researchers tested their method on the RAVDESS dataset (vocal expressions and songs). One of the most striking findings was that AlexNet, a relatively shallow model by modern standards, performed the best.
- AlexNet: 79.74% Accuracy
- ResNet-4: 75.89% Accuracy
- GoogleNet: 62.89% Accuracy
The authors hypothesize that because the pre-processing already "simplifies" the data by removing noise, hyper-deep networks like GoogleNet (100+ layers) may actually over-abstract the features, leading to poor performance.
Table 1: Comparison of the proposed model against human raters and other visual/acoustic baselines.
Critical Insight: More is Not Always Better
The study reveals a vital lesson in machine learning: Smart pre-processing can reduce the need for model complexity. By transforming the video into a "difference vision," the researchers effectively performed manual feature engineering that "taught" the AI to focus on muscle contraction rather than identity or lighting.
Limitations
- Pose Sensitivity: The current method relies heavily on frontal face alignment. If a subject turns their head significantly, the subtraction becomes chaotic.
- Real-time Constraints: While the CNN inference is fast, the alignment and subtraction steps currently limit smooth live-video analysis without further optimization.
Conclusion
This work demonstrates that "seeing the change" is more efficient than "seeing the whole." By achieving a 4.78% improvement over human accuracy using only visual data, the subtraction pre-processing method provides a lightweight yet powerful alternative to complex recurrent architectures in the field of affective computing.
