From YouTube to the Boardroom: Can Vlogs Predict Your Meeting Persona?
Cross-Domain Personality Prediction: From Video Blogs to Small Group Meetings
This paper explores cross-domain personality prediction, specifically targeting the "Extraversion" trait. It proposes a transfer learning approach that leverages a large-scale source domain of YouTube video blogs (vlogs) to improve prediction accuracy in a target domain of small-group physical meetings, achieving a peak accuracy of 70.4%.
Executive Summary
TL;DR: This study pioneers the use of social media data (YouTube vlogs) as a source domain to predict personality traits in physical small-group meetings. By applying domain adaptation techniques to visual nonverbal cues, the researchers demonstrated that a model trained on vloggers can predict the "Extraversion" of meeting participants with over 70% accuracy, significantly outperforming models trained purely on meeting data.
Background Positioning: Published in ICMI, this work acts as a bridge between Social Signal Processing (SSP) and Transfer Learning. It addresses the "small data" problem in behavioral science by tapping into the vast, albeit different, world of social media.
Problem & Motivation: The Context Collapse
In psychological research, data is a bottleneck. Recording and annotating high-quality videos of people interacting in small groups is expensive and time-consuming. Consequently, most datasets like the ELEA (Emergent LEAder) corpus are relatively small (approx. 100 participants).
The authors' core Insight is that video blogging is a form of "simulated conversation." When a vlogger speaks to a camera, they exhibit nonverbal cues—gestures, posture, and energy—similar to face-to-face interactions. However, the domains are not identical:
- Vlogs: Monologues, person is always the focus, high extraversion bias.
- Meetings: Interactions, turn-taking, more "neutral" personality distribution.
Methodology: Bridging the Gap
The study focuses on Extraversion, the trait most visibly encoded in body language.
1. Visual Feature Extraction
To maintain robustness across different camera setups, the authors used weighted Motion Energy Images (wMEI). This method collapses temporal motion into a single spatial representation, highlighting where and how much a person moves.
Figure: The pipeline from nonverbal feature extraction to cross-domain prediction.
2. Domain Adaptation Strategies
The paper tests several ways to combine the Source (VLOG) and Target (ELEA) data:
- SRC/TRG: Only source or only target data.
- COMB: A simple union of both datasets.
- AUG (Feature Augmentation): Mapping features into a higher-dimensional space that separates "domain-specific" features from "general" personality features.
Experiments & Results
The researchers conducted 10-fold cross-validation, varying the amount of "adaptation data" available from the target domain.
Figure: Comparison of Ridge Regression and SVM performance across different adaptation methods.
Key Findings:
- The "Data is King" Effect: Simply combining the datasets (COMB) yielded the best result (70.4%).
- Social Media as a Strong Prior: Even with zero target training data (the SRC approach), the vlog-trained model achieved 68% accuracy. This suggests that the visual expression of extraversion is remarkably consistent across these two very different social contexts.
- Baseline Defeat: The TRG-only model (the traditional approach) performed the worst, proving that the small size of meeting datasets limits the ability of a classifier to learn a robust representation of personality.
Critical Analysis & Conclusion
Takeaway
The study confirms that social media behavior is not "isolated" or "fake" in a way that makes it useless for real-world application; rather, vlogs are a rich repository of human social signals that can bootstrap models for more complex, low-data scenarios like corporate meetings or social psychology studies.
Limitations
- Trait Specificity: The study focused heavily on Extraversion. Traits like Conscientiousness or Agreeableness are far more "subtle" and may not transfer as easily using only motion-based visual cues.
- Modality: The authors intentionally ignored audio to focus on the domain gap, but in real meetings, vocal prosody and turn-taking dynamics are crucial indicators of personality.
Future Work
The next frontier involves Multimodal Domain Adaptation—combining text (speech-to-text), audio, and video—and exploring "unsupervised" adaptation where no labeled data from the target domain is available at all.
