Beyond Pixels: Landmark-Based Robustness in Emotion Recognition
An Adversarial Attacks Resistance-based Approach to Emotion Recognition from Images using Facial Landmarks
The paper introduces a facial landmark-based approach for Emotion Recognition (ER) designed specifically to resist adversarial attacks. By utilizing Dlib for landmark extraction and a Random Forest (RF) classifier, the method achieves robustness comparable to ResNet on clear data while significantly outperforming it under adversarial perturbations.
TL;DR
Deep learning models are notoriously fragile—a tiny bit of noise can flip a "Happy" prediction to "Angry." This paper demonstrates that by moving away from raw pixel-based classification and toward Facial Landmark Geometry, we can build emotion recognition systems that are virtually immune to common adversarial attacks (FGSM) and occlusions, while being 100x faster than traditional CNNs like ResNet.
Problem & Motivation: The "Homogeneous" Trap
Current SOTA models for emotion recognition often treat the task as a standard image classification problem. They learn the color distribution and pixel gradients (homogeneous knowledge) rather than the underlying anatomy of an expression.
The authors argue that this makes models susceptible to:
- Adversarial Attacks: Infinitesimal noise that shifts pixel values but doesn't change the expression to a human.
- Environmental Variance: Changes in skin color, lighting, or accessories (like glasses) that shouldn't affect emotional state analysis.
Inspired by Ekman’s Facial Action Coding Systems (FACS), the researchers hypothesize that the relative movement of facial landmarks is the true signal, and everything else is just noise.
Methodology: The Power of Relative Geometry
The core of the proposed method is to strip away the "distraction" of pixels.
1. Feature Extraction
Using the Dlib library, the team extracts (x, y) coordinates for key facial features (eyes, eyebrows, nose, mouth, jaw).
2. Spatial Normalization (The Central Point)
Directly feeding raw coordinates into a classifier fails because people appear at different positions in a frame. To solve this, the authors compute a Central Point (C)—the average of all detected coordinates. They then calculate the Euclidean distance of every landmark relative to C.
Fig. 1: Normalizing landmarks by calculating distances relative to a central point.
3. Classification
Instead of a heavy ResNet, they use a Random Forest (RF). Since the feature space is now reduced to a streamlined vector of distances, the computational overhead vanishes.
Adversarial Testing: ResNet vs. Geometry
The authors subjected both a ResNet and their Landmark-RF model to three types of attacks:
- Type A/B: Physical occlusions (black boxes over eyes or jaw).
- FGSM (Fast Gradient Sign Method): Mathematical noise designed to maximize model loss.
Fig. 2: FGSM attack on ResNet. The human sees no difference, but the model's confidence collapses or predicts the wrong emotion.
Experimental Results
The findings were stark. On the CK+ dataset:
- Clean Data: Both ResNet and Landmarks achieved ~97% accuracy.
- Under Attack: ResNet’s accuracy plummeted to 80.8% (Type A) and 90.8% (Type B).
- The Proposed Method: Maintained an astonishing 96.8% - 97% accuracy, showing almost zero sensitivity to the noise.
| Attack Type | ResNet Accuracy | Proposed (Landmark) Accuracy |
|---|---|---|
| None | 97.43% | 97.14% |
| Type A (Occlusion) | 80.86% | 96.86% |
| FGSM (Noise) | 90.00% | 98.00% |
Efficiency Gains
Perhaps the most "product-ready" takeaway: while ResNet took over 5 hours to process the data on a GTX 1080ti, the landmark approach finished in under 2 minutes.
Critical Analysis: Why This Works
The "Secret Sauce" lies in the information bottleneck. By forcing the model to only "see" landmark coordinates, the authors have effectively:
- Filtered the Noise: FGSM targets pixel gradients. If the model doesn't look at pixels, the attack has no vector.
- Invariance to Color: Histograms (as seen in the paper's Fig. 7) show that while occlusions destroy color distribution, the geometric pattern of a surprise or disgust expression remains intact.
Conclusion & Future Work
This paper is a strong reminder that "Deep Learning" isn't always the only answer. By incorporating domain-specific structural knowledge (Facial Landmarks), we can achieve Adversarial Robustness that pure pixel-crunching models currently lack.
The next hurdle? Testing this on "in-the-wild" datasets where landmark detection itself might be compromised by extreme head poses. If the landmark extractor falls, the system falls—making the robustness of the extractor itself the new frontier for research.
