Personality Prediction System: Deep Learning vs. Traditional ML on Facebook Data

Personality Prediction System from Facebook Users

2017-01-01
Tommy Tandera, Hendro, Derwin Suhartono, Rini Wongso, Yen Lina Prasetio
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a Personality Prediction System that analyzes Facebook status updates using the Big Five Model (Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism). It compares traditional Machine Learning (SVM, LDA, etc.) with Deep Learning architectures (MLP, LSTM, CNN), ultimately achieving a superior average accuracy of 74.17%.

    ## TL;DR
    This research investigates the effectiveness of predicting human personality traits from Facebook status updates. By comparing traditional machine learning (ML) with modern deep learning (DL) architectures, the study demonstrates that DL—specifically Multi-Layer Perceptrons (MLP) and hybrid LSTM-CNN models—outperforms legacy methods, reaching accuracies as high as 93.33% for specific traits.

    ## Executive Summary
    In the digital age, our "digital footprints" are mirrors of our psyche. This paper positions itself at the intersection of **Natural Language Processing (NLP)** and **Psychometrics**. Utilizing the **Big Five Model Personality** (Ocean: Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism), the authors provide a comprehensive analysis of how different algorithmic scenarios (feature selection, resampling, and architecture choice) impact the accuracy of behavioral prediction.

    ## The Problem: Why Questionnaires Are Not Enough
    Traditional personality recognition requires active user participation in lengthy surveys. In the context of "Big Data," this is unscalable. While early automated systems used **LIWC (Linguistic Inquiry and Word Count)** to find correlations between word categories and traits, these "closed-vocabulary" methods are limited by predefined dictionaries. The challenge lies in moving toward "open-vocabulary" methods that allow neural networks to discover latent patterns in how people express themselves.

    ## Methodology: Two Paths to the Subconscious

    ### 1. Traditional Machine Learning (The Baseline)
    The authors tested five core algorithms: **Naive Bayes, SVM, Logistic Regression, Gradient Boosting, and LDA**. 
    - **Features**: LIWC2015, SPLICE (Linguistic cues), and SNA (Social Network Analysis).
    - **Validation**: 10-fold cross-validation.

    ### 2. Deep Learning (The Innovation)
    The system shifts from manual feature engineering to **Word Embeddings (GloVe)**. 
    - **Architectures**: MLP, LSTM, GRU, and a custom **Hybrid LSTM+CNN 1D**.
    - **Preprocessing**: Removal of URLs, symbols, and stop words, alongside stemming to reduce dimensionality.

    ![Model Scenarios Table](https://cdn.atominnolab.com/wisdoc/tables/20260527-76aa5c73-96ba-41ba-96cf-1a74fbb3c0b7/page_003_block_009.png)
    *Table 3: The exhaustive list of experimental scenarios combining features, selection methods, and resampling.*

    ## Key Insights: Does "More Data" Mean "More Accuracy"?
    The researchers uncovered several technical nuances:
    - **The Resampling Effect**: Personality data is naturally imbalanced (e.g., more "Open" individuals in social media samples). **Under-sampling** significantly improved DL performance by preventing the model from becoming biased toward the majority class.
    - **The Architecture Winner**: While there was no "one-size-fits-all" architecture, **MLP** dominated the myPersonality dataset, while the **LSTM+CNN 1D** hybrid was exceptionally effective for a manual dataset, particularly for "Extraversion."

    ![ML Results](https://cdn.atominnolab.com/wisdoc/tables/20260527-76aa5c73-96ba-41ba-96cf-1a74fbb3c0b7/page_005_block_003.png)
    *Table 4: Performance of traditional ML—LDA achieved the highest average across traits (63.04%).*

    ![DL Results](https://cdn.atominnolab.com/wisdoc/tables/20260527-76aa5c73-96ba-41ba-96cf-1a74fbb3c0b7/page_005_block_007.png)
    *Table 6: Deep Learning results—MLP surged ahead with a 70.78% average, showcasing the power of neural modeling.*

    ## Critical Analysis & Future Outlook
    While the study shows a clear victory for Deep Learning, it acknowledges a common bottleneck: **Dataset Size**. With only 250-400 users, the "Deep" in Deep Learning is somewhat constrained. 

    **Takeaways for the Industry:**
    1. **Hybrid Models Work**: Combining the temporal sensitivity of LSTM with the spatial pattern recognition of CNN yields high precision in linguistic analysis.
    2. **Preprocessing is Critical**: For non-English data (like the Bahasa samples used), manual slang replacement and careful translation are prerequisites for maintaining the semantic integrity of personality cues.

    **Limitations**: The reliance on "Apply Magic Sauce" for labeling the manual dataset introduces a third-party dependency that might influence the ground truth. Future work using XGBoost and larger, natively labeled datasets could push these accuracies even closer to human-level (or beyond) judgment.

    ## Conclusion
    This research successfully demonstrates that by moving away from "count-based" linguistic features toward "context-aware" neural embeddings, we can predict human personality with high accuracy. The 74.17% average accuracy sets a strong precedent for using AI as a tool for large-scale behavioral research and personalized user experiences.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Transformer-based models like BERT or RoBERTa for Big Five personality prediction from social media text.
  • Which study first introduced the "myPersonality" dataset, and how have subsequent benchmarks evolved in terms of accuracy for the 'Neuroticism' trait?
  • Explore how multi-modal deep learning, combining text, images, and social network graphs, has been applied to personality profiling in recent psychological research.
Contents
Personality Prediction System: Deep Learning vs. Traditional ML on Facebook Data
1. TL;DR
2. Executive Summary
3. The Problem: Why Questionnaires Are Not Enough
4. Methodology: Two Paths to the Subconscious
4.1. 1. Traditional Machine Learning (The Baseline)
4.2. 2. Deep Learning (The Innovation)
5. Key Insights: Does "More Data" Mean "More Accuracy"?
6. Critical Analysis & Future Outlook
7. Conclusion