Deciphering First Impressions: Can AI Predict Personality Using Only Video?

Single-Modal Video Analysis of Personality Traits using Low-Level Visual Features

2020-11-09
Daniel Helm, Martin Kampel
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a single-modal video analysis framework for predicting Big-Five personality traits using only low-level visual features. By evaluating standard 2D-CNN and 3D-CNN architectures, the authors achieve a competitive Mean Average Accuracy of 0.8905 on the ChaLearn First Impressions dataset without relying on audio modalities.

TL;DR

First impressions are formed in seconds, and while previous research combined voice and sight, this paper proves that visual low-level features alone can predict "Big-Five" personality traits with nearly 90% accuracy. Using 3D-CNNs and refined face-tracking, the authors demonstrate that how you move and present yourself on camera speaks volumes—even when the sound is off.

Problem & Motivation: The Multi-Modal Dependency

In the quest to automate human personality analysis—useful for everything from job interview coaching to better social robots—the current SOTA (State-of-the-Art) has a "crutch": multi-modal fusion. Most winning solutions from challenges like ChaLearn LAP rely heavily on audio features (volume, pitch, rate) to bolster visual data.

The authors identify a critical gap: To what extent does the visual signal stand on its own? Furthermore, training a model to output five continuous values (regression) between 0 and 1 is significantly more complex than simple classification. The challenge lies in capturing both the static "look" of a person and the dynamic "feel" of their movement without the noise of the background.

Methodology: The Core Architectures

The paper explores two distinct technological paths:

  1. Approach 1: The Image-Based Baseline: Adaptive VGG-16 models that process individual frames and average the results.
  2. Approach 2: The Temporal Specialist: A 3D-CNN that processes video volumes, capturing the "flow" of expressions over time.

The Pre-processing Pipeline

The authors highlight that raw video is too noisy. They implemented a pipeline using Dlib to detect 68 facial landmarks, allowing them to rotate and crop faces so that eyes are always horizontal—ensuring the "low-level features" are standardized across 10,000 YouTube videos.

Face Extraction Pipeline

Interestingly, the researchers found that "Extended-Image-Features" (which include the hair, shoulders, and a bit of background) provided better context for traits like Openness and Extraversion than a tight face-only crop.

Architecture Comparison

Experiments & Results: 3D-CNN Wins

The results show a clear hierarchy in performance. While the 2D-CNN reached an accuracy of 0.8869, the 3D-CNN (Approach 2) took the lead with 0.8905.

Key Insights from the Data:

  • Loss Functions Matter: While MSE is the standard for regression, the authors found that Binary Cross-Entropy (BC) combined with a Sigmoid activation layer often yielded more stable results when predicting traits in the 0-1 range.
  • The Overfitting Trap: Without data augmentation and Weight Decay (L2 regularization), the models quickly "memorized" the training set, leading to poor validation performance (as seen in the fluctuating loss curves).
  • Distribution Alignment: The 3D-CNN produced a standard deviation in traits that more closely matched human ground-truth labels compared to the "flatter" 2D model.

Training History Analysis

Critical Analysis & Conclusion

This work provides a fundamental baseline for single-modal analysis. By stripping away audio, it proves that "low-level" visual cues (the geometry of a smile, the tilt of a head) are powerful predictors of perceived personality.

Takeaways & Limitations

  • The "Why" vs. the "What": Through feature map analysis (Deconvolution), the authors noticed the models sometimes focused on less relevant image areas. This suggests the "black box" of personality AI still needs more fine-tuning to ensure it's looking at human behavior rather than background artifacts.
  • Ethical Warnings: The authors conclude with a sobering reminder: while these tools can help people practice for interviews, they also carry the risk of discrimination. If an algorithm decides your "Conscientiousness" score is low before you even speak, the potential for misuse in HR is high.

Future Work

The next frontier involves exploring High-Level features (like specific gesture patterns) and addressing the ethical bias inherent in these automated systems to ensure they don't just mimic human prejudices.


The source code for this project is available on GitHub for further academic exploration.

Find Similar Papers

Try Our Examples

  • Find recent papers on single-modal personality trait recognition that focus specifically on hand gestures and body language versus facial expressions.
  • Which study first introduced the ChaLearn First Impressions dataset, and what were the primary baselines established for visual-only regression?
  • Explore research that applies 3D-CNN or SlowFast networks to socio-psychological trait analysis in Human-Computer Interaction.
Contents
Deciphering First Impressions: Can AI Predict Personality Using Only Video?
1. TL;DR
2. Problem & Motivation: The Multi-Modal Dependency
3. Methodology: The Core Architectures
3.1. The Pre-processing Pipeline
4. Experiments & Results: 3D-CNN Wins
4.1. Key Insights from the Data:
5. Critical Analysis & Conclusion
5.1. Takeaways & Limitations
5.2. Future Work