Hybrid Intelligence: Combining Random Forest and SIR for Weibo Diffusion Prediction

Modeling of Information Diffusion in Sina Weibo Based on Random Forest Classifier and SIR Model

2019-11-06
Zhang Jianyi, He Ping, Ken K. T. Tsang, Deng Yuhui
Summary
Problem
Method
Results
Takeaways
Abstract

This paper proposes a hybrid framework for predicting information diffusion in Sina Weibo by integrating a Random Forest classifier with the Susceptible-Infected-Recovered (SIR) epidemic model. By utilizing 15 node and edge features, the model predicts individual reposting probabilities to drive macroscopic diffusion simulations.

TL;DR

Predicting how a "hot topic" spreads across social media is notoriously difficult due to the complexity of human behavior and network topology. This paper introduces a dual-layered approach: using a Random Forest classifier to predict if an individual will repost a message, and then feeding those probabilities into a SIR (Susceptible-Infected-Recovered) epidemic model to forecast the overall lifecycle of a microblog.

Problem & Motivation: The Limits of Average Probability

Traditional epidemiological models used for information spread often rely on a constant "infection rate." In the context of Sina Weibo, this incorrectly assumes that every user has the same likelihood of sharing a post.

The authors identify two critical gaps:

  1. Feature Neglect: Existing models ignore node-specific features like historical interaction frequency, gender, and interest similarity.
  2. Data Imbalance: In reality, users ignore 90% of what they see. Standard models trained on such data tend to predict "not repost" for everything to achieve high nominal accuracy, failing to capture the actual "viral" moments.

Methodology: From Individual Behavior to Global Trends

1. Feature Engineering & Balancing

The researchers identified 15 key features across three dimensions: blogger influence, fan characteristics, and content properties. To solve the data skewness, they employed SMOTE (Synthetic Minority Over-sampling Technique), which creates synthetic examples of the minority "repost" class to ensure the classifier learns what actually triggers a share.

2. The Random Forest Engine

Why Random Forest? It handles categorical data and non-linear relationships better than traditional regressions. By analyzing "Gini Importance," the study found that Historical Forwarding Frequency and Interest Similarity were the most potent predictors of a repost.

Table 1: 15 Features of Nodes and Edges

3. SIR Simulation

The output of the Random Forest—the probability of a user reposting—becomes the (infection rate) in the SIR differential equations:

  • S (Susceptible): Users exposed to the microblog.
  • I (Infected): Users who repost the message.
  • R (Recovered): Users who have already reposted and moved on.

Experiments & Results

The hybrid model was validated against 2018 Sina Weibo data. The results showed that the Random Forest outperformed Support Vector Machines (SVM), particularly when combined with random under-sampling of the majority class.

Simulation Result vs Actual Data The figure shows the time history of the "Recovered" population (black curve) closely matching the actual accumulated reposts (blue curve) over time.

Key Performance Metrics:

  • Classification Precision: >80%.
  • Simulation Error: <15% deviation from real-world time-series data.
  • Top Feature: "Historical forwarding frequency" (0.26 importance score) proved that past interaction is the strongest indicator of future engagement.

Critical Analysis & Conclusion

Takeaway

The core insight of this work is that micro-level predictions drive macro-level accuracy. By moving away from "average probability" and toward "calculated probability" per node, the SIR model becomes a powerful tool for public opinion monitoring.

Limitations & Future Work

  • Static Topology: The model assumes a fixed network structure, whereas real social networks are dynamic.
  • Feature Dimensionality: While 15 features are effective, the authors suggest that semi-supervised learning could further refine the model when labeled data is scarce.
  • Potential Extensions: Applying this framework to multi-modal data (images/videos) or cross-platform diffusion (Weibo to WeChat) remains an open research frontier.

Find Similar Papers

Try Our Examples

  • Find recent studies that integrate deep learning architectures, such as Graph Neural Networks (GNNs), with the SIR model for social network information diffusion.
  • What are the primary theoretical differences in information propagation between the SIR model and the Independent Cascade (IC) model in the context of directed social graphs?
  • Search for research investigating how the SMOTE technique performs against Cost-Sensitive Learning in the specific domain of predicting rare user interactions in social media.
Contents
Hybrid Intelligence: Combining Random Forest and SIR for Weibo Diffusion Prediction
1. TL;DR
2. Problem & Motivation: The Limits of Average Probability
3. Methodology: From Individual Behavior to Global Trends
3.1. 1. Feature Engineering & Balancing
3.2. 2. The Random Forest Engine
3.3. 3. SIR Simulation
4. Experiments & Results
4.1. Key Performance Metrics:
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work