Retweet Prediction: Leveraging the Intersection of Topic, Emotion, and Personality

Retweet Prediction based on Topic, Emotion and Personality

2021-09-01
Syeda Nadia Firdaus, Chen Ding, Alireza Sadeghian
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a retweet prediction framework that integrates explicit content features with high-dimensional implicit behavioral features—namely Topic, Emotion, and Personality (TEP). The authors propose two distinct modeling paradigms: a classification approach using XGBoost/Random Forest and a Matrix Factorization (MF) approach with novel tweet-similarity regularization terms, achieving 5%-9% and 4%-6% F1-score improvements over respective baselines.

TL;DR

Retweeting is the heartbeat of Twitter's information diffusion. While early models focused on "what" was said (hashtags, keywords), this paper dives into "who" is sharing and "how" they feel. By merging Twitter-LDA topics, 10-dimensional emotion vectors, and 35 personality facets, the researchers built a prediction engine that outperforms traditional baselines by up to 9% in F1-score using both XGBoost and custom Matrix Factorization.

Problem & Motivation: Beyond Keywords

Why do we retweet? It’s rarely just about a hashtag. Existing research often hit a glass ceiling by ignoring the psychological "Inductive Bias" of the user. Most models utilized binary sentiment (Positive vs. Negative), but human emotion is a spectrum (Joy, Fear, Disgust, etc.). Furthermore, personality traits—like Neuroticism or Openness—dictate our propensity to spread information.

The authors observed a critical gap: Personality conflict or emotional divergence can stop a retweet even if the topic matches the user's interest. Their insight was to treat a user's past timeline as a psychological mirror to predict future "virality" susceptibility.

Methodology: The TEP Framework

The core of this work lies in the Profile Generator, which translates raw text into a multidimensional behavioral space.

1. The Feature Trio

  • Topic (Twitter-LDA): Unlike standard LDA, Twitter-LDA assumes one tweet has one topic, reducing noise in short-form text.
  • Emotion (Plutchik’s Wheel): Using NRC lexicons, the model builds a 10D vector (8 emotions + 2 sentiments).
  • Personality (Big Five): 35-dimensional scores covering traits and facets (e.g., Curiousness, Temperament) derived from Linguistic Inquiry and Word Count (LIWC) correlations.

2. Matrix Factorization with a Twist

While the classification approach (XGBoost) calculates similarity scores between user and tweet, the Matrix Factorization approach introduces a "Content-Aware" latent space.

Architecture of the retweet prediction model

The authors proposed a new regularization term: If two tweets are similar in the observed space (Topic/Emotion/Personality), they must remain close in the latent space. This effectively forces the model to learn that "similar people retweet similar-vibe content."

Experiments & Results

The researchers tested their models on 1.6 million posts from 1,136 users.

Key Breakthroughs:

  • Information Source: A major discovery was that retweet-only profiles are nearly as effective as "tweet+retweet" profiles but much faster to process. This suggests that the act of "curating" (retweeting) reveals more about our sharing behavior than the act of "creating" (tweeting).
  • Model Performance:
    • XGBoost (F-full): Highest F1-score, providing a 5-9% lift over TF-IDF and standard LDA baselines.
    • Matrix Factorization (Approach 2): Bested the basic MF baseline by 6% in F1-score, notably excelling in Recall.

XGBoost Performance Comparison

Precision vs. Recall Trade-off

An interesting academic takeaway: Machine Learning (Classification) provides higher Precision (avoiding false positives), making it ideal for Tweet Recommendation. However, Matrix Factorization provides higher Recall (finding all potential spreaders), making it the superior choice for Marketing Campaigns.

Deep Insight & Conclusion

This paper shifts the retweet prediction paradigm from "content-matching" to "psychological-profiling." The fact that Personality features consistently improved the Recall across all models proves that personality is a "subtle but strong" anchor for social behavior.

Limitations: The model is currently optimized for active users (100+ posts). For "lurkers" or less active accounts, the sparsity of emotional data remains a challenge.

Future Outlook: The next frontier is Behavioral Drift—how our personality stays static while our topical interests and emotional states shift over time. Integrating temporal dynamics into this TEP framework could unlock true real-time social sensing.

Find Similar Papers

Try Our Examples

  • Examine recent literature on the integration of Big Five personality traits and 10-dimensional emotion vectors for fake news detection or information diffusion modeling.
  • Trace the origin of the Twitter-LDA model and identify subsequent improvements in extracting single topics from short-form social media texts.
  • Search for matrix factorization studies that utilize content-based regularization terms to solve the cold-start problem in social media recommendation systems.
Contents
Retweet Prediction: Leveraging the Intersection of Topic, Emotion, and Personality
1. TL;DR
2. Problem & Motivation: Beyond Keywords
3. Methodology: The TEP Framework
3.1. 1. The Feature Trio
3.2. 2. Matrix Factorization with a Twist
4. Experiments & Results
4.1. Key Breakthroughs:
4.2. Precision vs. Recall Trade-off
5. Deep Insight & Conclusion