Reimagining Information Diffusion: When Social Networks Meet Recommender Systems
Recommender System-Based Diffusion Inferring for Open Social Networks
This paper introduces DIM-SPTF, a novel framework that casts the information diffusion inference task as a recommendation problem within Open Social Networks (OSNs). By leveraging a Recurrent Neural Network (RNN) to model temporal and social features, the method achieves superior performance in reconstructing diffusion networks compared to traditional cascade-based and non-cascade-based SOTA baselines.
TL;DR
Understanding how information spreads in Open Social Networks (OSNs) is a "holy grail" for community detection and network analysis. This paper introduces DIM-SPTF, a framework that treats information diffusion as a recommendation process. By feeding user social preferences and text features into an RNN-based recommender system, the authors achieve a ~10% accuracy boost in reconstructing hidden diffusion networks.
Background & Motivation: Why Current Models Fail
The diffusion of a tweet or a blog post is rarely just a function of time. Yet, many classical models (SOTA) rely heavily on simple timestamps or treat features (like content and user ID) as independent variables.
The authors identify two fatal flaws in prior work:
- Ignoring Social DNA: Users have distinct preferences (e.g., sports vs. politics) that dictate their likelihood to "propagate" a specific topic.
- Feature Isolation: In reality, text features and reward-seeking social behaviors are deeply intertwined; they do not perform independently in the diffusion process.
Methodology: The "Recommendation" Intuition
The core innovation of DIM-SPTF (Diffusion Inferring Method based on Recommender System) is its conceptual bridge: A social network user "infecting" another with information is essentially the system "recommending" a commodity (the info) to a consumer (the peer).
1. Feature Extraction (The Input)
Instead of just recording (Time, UserID), the authors expand the "Cascade" model to include:
- User Preference (): A 9-dimensional vector (Science, Arts, etc.) derived from historical behaviors using a mapping matrix.
- Text Features (): Semantic relevance of the post to the aforementioned nine categories, calculated using keyword extraction and semantic distance.
2. The RNN Recommender Engine
Since diffusion is a sequential event, the authors employ a Recurrent Neural Network (RNN). The RNN’s hidden state acts as the "memory" of the network, capturing the temporal correlations between adjacent users in a cascade.
Figure 1: The architecture of DIM-SPTF, illustrating the flow from user data to diffusion relationship judgment.
Experimental Battleground
The researchers tested DIM-SPTF against DDNE (a dynamic cascade model) and FNI (a feature-based model) using both synthetic data and real-world Sina Microblog data.
Performance Gains
- Accuracy: DIM-SPTF consistently led by ~10% over FNI.
- Precision: In larger datasets, the precision reached 61%-78%, outperforming others by significant margins.
- Robustness: One of the most impressive results was the model's resistance to time delay. While DDNE and FNI saw performance plummets as delay increased (up to 5 hours), DIM-SPTF remained relatively stable.
Figure 2: Accuracy and Precision metrics across different data scales—DIM-SPTF shows a clear lead.
Visualizing the "Hidden" Network
The authors visualized the reconstructed network of the "USA Huawei ban" topic. The results were telling:
- Self-Media dominance: While official media nodes have high influence and speed, self-media users make up the vast majority of the "diffusion body."
- Preference Clustering: The inferred graph showed users focusing on politics, technology, and economy—perfectly matching the topic's nature.
Figure 3: A visualization of the inferred diffusion network, showing node influence and clustering.
Critical Insight & Conclusion
The success of DIM-SPTF suggests that social context is not just "extra credit"—it is foundational. By quantifying user preferences as a static "Social DNA" and information as "Commodities," we can move away from pure probabilistic modeling toward intent-based modeling.
Limitations: The model assumes user preferences are static during the diffusion process. In reality, a major event might shift a user's interests overnight. Future work integrating dynamic topology theory and shifting social relationships could further sharpen the accuracy of these "digital mirrors" of human interaction.
