personality2vec: Fusing Social Structure and Linguistic Cues for Better Personality Prediction

Personality2vec: Network Representation Learning for Personality

2020-07-01
Zhanming Guan, Bin Wu, Bai Wang, Hezi Liu
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces personality2vec, a novel Network Representation Learning (NRL) framework designed to predict Big Five personality traits from online social network (OSN) texts. It achieves State-of-the-Art (SOTA) performance across three major datasets (MyPersonality, YouTube, PAN2015) by integrating semantic, linguistic, and structural information into a unified user embedding.

TL;DR

Predicting personality from social media is notoriously difficult due to "noisy" text and small datasets. personality2vec solves this by treating users as nodes in a graph. By combining what users say (Linguistic/Semantic similarity) with where they sit in a network (Structural information), it generates high-quality "Personality Vectors." The model outperforms traditional deep learning and dictionary-based methods across various benchmarks.

Background: Why Text Isn't Enough

Most existing personality models act like "lone readers"—they look at a user's posts in a vacuum. However, personality is inherently social. Previous attempts faced a "Small Data" wall: deep learning models require massive amounts of labeled data, but professional psychological labeling is expensive and subjective.

The authors' core insight is that if two users exhibit similar linguistic patterns and occupy similar structural positions in a similarity network, they likely share personality traits. By using Network Representation Learning (NRL), they can "borrow" information from similar nodes to help predict the scores of others.

Methodology: The Three Pillars of Personality2vec

The framework consists of three innovative stages:

1. Multi-Dimensional Network Construction

The model builds a graph where edges aren't just "follows," but represent Text Similarity. This similarity is a product of:

  • Semantic Info: Calculated via TF-IDF.
  • Linguistic Info: Calculated using the LIWC dictionary and 10 custom "OSN Special Features" (like the rate of consecutive exclamation marks or capital letters).

2. Biased Walk Strategy

Instead of a random walk, the model uses a "biased walk" (influenced by node2vec). It prefers nodes that are either structurally important (high degree) or linguistically similar. This ensures the "corpus" generated for the language model contains sequences of "kindred spirits."

Overall framework

3. Personality-Aware Skip-gram

This is where the magic happens. The authors modified the standard Word2Vec/Skip-gram approach in two ways:

  • New Huffman Tree: Traditional trees cluster words by frequency. Here, nodes are grouped by Path Frequency, ensuring that nodes with similar personality features are close together in the hierarchy, preventing "gradient noise" from unrelated nodes.
  • Label-Informed Gradient: They introduced a sim(w, u) term into the loss function. If two users have known similar personality scores, the model pushes their vectors closer together even more aggressively during back-propagation.

New Huffman Tree Construction

Experiments: Superior Consistency

The researchers tested the model on MyPersonality (Facebook), YouTube, and PAN2015.

Key Findings:

  • Deep Learning Failure: Models like AttRCNN-CNN struggled because the datasets (only a few hundred to a few thousand users) were too small for heavy neural weights.
  • Structural Advantage: NRL-based methods (DeepWalk, node2vec) performed better, but personality2vec’s addition of Linguistic Sentiment gave it the winning edge.
  • MAE Reduction: The model achieved significant improvements in Mean Absolute Error (MAE) across all Big Five traits (Openness, Conscientiousness, Extroversion, Agreeableness, Neuroticism).

Experimental Results Table

Critical Insight & Conclusion

The true value of personality2vec lies in its robustness to small data. By transforming a text classification problem into a graph representation problem, it effectively handles the sparse and noisy nature of social media language.

Limitations: The model still relies on a threshold-based graph construction, which might be sensitive to the quality of the initial semantic/linguistic extraction.

Future Work: The authors suggest integrating Knowledge Graphs to bring professional psychological knowledge directly into the embedding process—a promising step toward "Physiological-AI" fusion.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Neural Networks (GNNs) or Graph Transformers specifically for personality trait prediction in social media.
  • Which paper originally proposed the "Adawalk" kernel for personality detection, and how does the personality-weighted skip-gram in personality2vec differ from its weight update mechanism?
  • Explore how the Big Five personality traits are being integrated into personalized recommendation systems using network representation learning techniques.
Contents
personality2vec: Fusing Social Structure and Linguistic Cues for Better Personality Prediction
1. TL;DR
2. Background: Why Text Isn't Enough
3. Methodology: The Three Pillars of Personality2vec
3.1. 1. Multi-Dimensional Network Construction
3.2. 2. Biased Walk Strategy
3.3. 3. Personality-Aware Skip-gram
4. Experiments: Superior Consistency
5. Critical Insight & Conclusion