Decoding Emotion through Motion: A Silhouette-Based Approach to Affective Computing
Emotion Evaluation Through Body Movements Based on Silhouette Extraction
This paper presents a non-intrusive emotional evaluation framework that utilizes OpenCV for body silhouette extraction and the PAD (Pleasure, Arousal, Dominance) scale for real-time emotional ground-truth mapping. By analyzing temporal behavioral features in response to Chinese folk music induction, the system establishes a robust Emotion-Behavior Library capable of distinguishing basic emotional states.
TL;DR
Researchers have developed a non-invasive method to evaluate human emotions by analyzing the "geometry of movement." By extracting body silhouettes from video and mapping them against the PAD (Pleasure, Arousal, Dominance) emotional scale, the system can distinguish between positive and negative emotional states with 72.5% accuracy, utilizing culturally specific music to induce natural physical responses.
Background & Motivation
Affective Computing aims to bridge the gap between human emotion and machine logic. While facial expressions and speech analysis are common, they are often prone to "acting" or external noise. Body movement, however, often provides a more truthful, subconscious reflection of our internal state.
The authors identify a critical gap: most existing behavioral recognition systems rely on static postures or invasive wearable sensors. This study seeks to answer: Can we objectively quantify emotion using only the dynamics of a person's silhouette?
Methodology: From Pixels to Psychology
The workflow involves three critical stages: Emotional Induction, Feature Extraction, and Mapping.
1. The Stimulus: Chinese Folk Music
Unlike previous studies that favored Western classical music, this research uses a dedicated Chinese Folk Music Emotional Library. Music is used as the catalyst because it is immersive and provides a sustained emotional state suitable for real-time behavioral observation.
2. Feature Extraction via OpenCV
The system doesn't track specific joints. Instead, it looks at the Silhouette () and the Minimum External Polygon (). Two key mathematical metrics are defined:
- CI [t] (Compactness Index): The ratio of silhouette area to its bounding polygon, representing how "spread out" or "contained" a movement is.
- RoSC (Rate of Silhouette Change): A measure of the velocity of bodily expansion or contraction.
Fig 1. The systematic flow from music induction to silhouette processing and final classification.
Experimental Insights & SOTA Contrast
The study highlights a significant difference in how we perceive and express emotions based on viewing angles.
- The Power of Perspective: The Front Camera was significantly more effective than the Right Camera. In binary classification (Positive vs. Negative), the front view achieved 72.5%, whereas the right view struggled at 52.5%, nearly equivalent to random guessing.
- Energy Signatures: The data shows that positive emotions correlate with higher "behavioral energy." As seen in the figure below, the amplitude of movement (Red line) is much larger for positive states than for negative ones (Blue line).
Fig 2. Velocity and acceleration of movement: Red (Positive) vs. Blue (Negative). The positive state demonstrates higher frequency and amplitude.
Results Summary Table
| Classification Type | Camera Angle | Accuracy (%) |
|---|---|---|
| Positive vs. Negative | Front | 72.50% |
| Positive, Neutral, Negative | Front | 50.00% |
| Positive vs. Negative | Right | 52.50% |
Deep Insight & Conclusion
This research confirms that body silhouettes are an "honest signal" of emotion. The primary takeaway is that positive emotions are characterized by "expansive and high-velocity" movements, while negative emotions are more "withdrawn and low-energy," closely resembling neutral states in a silhouette profile.
Limitations & Future Work: The drop in accuracy when moving from 2-way to 3-way classification suggests that "Neutral" and "Negative" emotions are behaviorally similar in a silhouette-only model. Future iterations could benefit from Multi-modal Fusion—combining silhouette dynamics with skeletal tracking or thermal imaging—to further refine the distinction between subtle affective states.
Ultimately, this work paves the way for more naturalistic HCI environments, such as "smart rooms" or "affective kiosks" that understand your mood simply by how you walk and move within them.
