Beyond Fixed Constants: Decoding Information Cascades via Personalized Retweet Probability
Information Propagation with Retweet Probability on Online Social Network
This paper introduces an enhanced information propagation framework that replaces homogenous retweet probabilities with personalized probabilities calculated via Logistic Regression (LR). By integrating these individual probabilities into Independent Cascade (IC) and Susceptible-Infectious-Susceptible (SIS) models, the authors achieve a more realistic simulation of information flow on the SINA Weibo platform.
TL;DR
Information propagation on social networks is often modeled like a biological virus. However, unlike a virus, information doesn't "infect" everyone with the same probability. This paper demonstrates that by replacing fixed average probabilities with individualized retweet probabilities learned from social features, we can predict information flow more accurately and discover that information actually spreads much faster than traditional models suggest.
Context: The Flaw in the "Dumb" Virus Analogy
Classic models like Independent Cascade (IC) and SIS (Susceptible-Infectious-Susceptible) treat users as uniform nodes in a graph. In these frameworks, if User A posts a message, User B retweets it based on a global constant .
The authors argue this is fundamentally flawed. In the real world, your decision to retweet depends on:
- Who you are (Are you a frequent poster? Are you verified?)
- Where you are (Geographic similarity)
- Your relationship (Do you follow each other mutually? Do you share many common friends?)
By ignoring these features, previous research has been "blind" to the true speed and scale of social media dynamics.
Methodology: Engineering the Probability
The core contribution is the shift from (a constant) to (a function of user 's context).
1. Feature Extraction
The researchers crawled SINA Weibo, extracting a dataset of over 21,000 users. They focused on two categories of features:
- User Info: Gender, Verified status, Tweet count, and Location similarity.
- Structure Information: Co-follow numbers and Mutual follow status (bidirectional edges).
2. Logistic Regression (LR) Model
They used a Logistic Regression model to define the retweet probability: By training on actual retweet history, the model learns the "weight" of each social feature. Interestingly, the authors found that Individual Models (trained specifically on a user's local neighborhood) significantly outperformed Global Models, emphasizing that retweeting is a highly contextual, local behavior.
Table 3: Individual LR models show higher F-values than Global models, justifying the personalized approach.
Simulation: The Speed of Truth (and Rumors)
The authors implemented the SIS-p and IC-p models. In the SIS-p model, information flow is treated as a directed process where a tweet "infects" a follower's timeline.
Key Discovery: Speed Acceleration
When using the learned probabilities, the simulation reached its peak much faster than the constant-probability version. In the IC model, the "personalized" version reached steady-state in 4 steps, while the "homogenous" version took 6 steps.
Fig 1 & 2: Notice how the red lines (Individual Probability) ramp up much steeper than the green lines (Constant Probability).
The "Zhang Yimou" Effect
The paper also highlights the importance of the Initial Poster. By comparing a common user ("wwwyyyddd") with a celebrity (Zhang Yimou), the authors show that high-influence nodes trigger a "vertical" jump in retweets within just 1-2 steps, whereas common users rely on "stochastic luck" to go viral.
Fig 3: The exponential gap between an influencer and a standard user in the IC-p model.
Critical Insight & Conclusion
The biggest takeaway is that homogenous models underestimate the "explosiveness" of social networks.
If we are trying to stop a rumor or an epidemic of misinformation, relying on old constant-probability models might give us a false sense of security, suggesting we have more time to react than we actually do. This research provides a roadmap for more surgical interventions: by calculating , platforms can identify precisely which "edges" in the network are the most dangerous conduits for rapid spread.
Limitations: The model currently lacks "content analysis" (the actual text of the tweet) and suffers from data sparsity for new users. Future work incorporating NLP (Natural Language Processing) could bridge this gap.
