Textual-Homo-IC: Content-Aware Information Diffusion in Dynamic Social Networks

Homophily Independent Cascade Diffusion Model Based on Textual Information

2018-01-01
Thi Kim Thoa Ho, Quang Vu Bui, Marc Bui
Summary
Problem
Method
Results
Takeaways
Abstract

The paper proposes Textual-Homo-IC, an expansion of the Independent Cascade (IC) model that estimates infection probabilities based on homophily derived from textual content. By integrating Latent Dirichlet Allocation (LDA) and the Author-Topic Model (ATM), the authors personalize diffusion probabilities on both static and dynamic agent-based networks.

TL;DR

Researchers have moved beyond simple "random" link probabilities in social spreading. The Textual-Homo-IC model leverages topic modeling (LDA/ATM) to calculate infection probabilities based on the actual semantic similarity (homophily) between users. Tested on NIPS and Twitter data, this method increases simulated influence reach by over 50% Compared to traditional Independent Cascade models.

Problem & Motivation: Beyond Random Probability

In traditional information diffusion models like the Independent Cascade (IC), we often assume that if User A is connected to User B, there is a fixed or random probability (e.g., a uniform distribution) that A will "infect" B with a piece of information.

However, human behavior is rarely random. We are much more likely to adopt an idea or share a paper if we share professional interests or social values with the source—a concept known as homophily. The researchers identified two major gaps in current literature:

  1. Semantic Blindness: Most models ignore the rich textual data (tweets, papers, bios) that define user interests.
  2. Static Property Fallacy: Most "dynamic" models only change the links between people, whereas in reality, our interests (node properties) evolve as we interact with others.

Methodology: Mapping Interests to Math

The core innovation lies in the Textual-Homo-IC framework, which operates in three stages:

1. Topic Distribution Estimation

The model treats each agent as a distribution of topics. For instance, a scientist might be 60% "Machine Learning," 30% "Optimization," and 10% "Ethics."

  • LDA (Latent Dirichlet Allocation): Used for general document-to-topic mapping.
  • ATM (Author-Topic Model): A more robust choice for authorship, linking multiple documents to a single agent's profile.

2. Homophily-Based Distance

Instead of a coin flip, the probability of infection is the inverse of the distance between their topic vectors. The authors utilized:

  • Hellinger Distance: Measures the overlap between two probability distributions.
  • Jensen-Shannon Divergence: A symmetric measure of similarity.

The final probability is defined as: .

Model Architecture Note: Above represents the conceptual flow of the diffusion process across agent nodes.

3. The Dynamic Social Loop

Unlike static models, Algorithm 3 allows the network to "evolve." As agents interact, their topic distributions are updated using an EM (Expectation-Maximization) iteration. This simulates a user "learning" or shifting interests over time, which in turn recalculates the infection probabilities for the next time step.

Experiments & Performance

The researchers benchmarked their model against Random-IC using two distinct datasets:

  1. NIPS Co-author Network: 2,479 scientists and 1,740 papers.
  2. Twitter: 1,524 users and their "follow" relations, including historical tweets.

Key Results:

  • Reach: On the NIPS dataset, using LDA + Hellinger distance resulted in 1369 active nodes, compared to only ~700 for the random baseline.
  • Saturation: In static networks, diffusion usually hits a "steady state" and stops. In the dynamic version of Textual-Homo-IC, the active number continues to fluctuate and grow as the network's internal properties shift.

Performance Comparison Fig: Comparison of diffusion reach across different configurations. Textual-Homo-IC consistently outperforms the baseline.

Critical Insight & Conclusion

This paper effectively argues that homophily is the engine of diffusion. By grounding the "infection probability" in actual textual data, the model provides a more realistic simulation of how ideas propagate in specialized communities (like academia) or interest-based platforms (like Twitter).

Limitations: While LDA and ATM were SOTA at the time of the paper's focus, they are "bag-of-words" models. They miss the nuanced context that modern Transformers (like BERT or Llama) could provide.

Future Outlook: The next generation of these models will likely replace LDA with Embedding distances from LLMs, allowing for an even more granular understanding of why certain ideas go viral while others die in the metadata.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Large Language Models (LLMs) instead of LDA for homophily-based information diffusion in social networks.
  • Who first proposed the Author-Topic Model (ATM), and how does the online EM update process compare to stochastic variational inference for scaling to millions of users?
  • Investigate studies that apply homophily-driven Independent Cascade models to epidemic modeling or the spread of misinformation in polarized digital environments.
Contents
Textual-Homo-IC: Content-Aware Information Diffusion in Dynamic Social Networks
1. TL;DR
2. Problem & Motivation: Beyond Random Probability
3. Methodology: Mapping Interests to Math
3.1. 1. Topic Distribution Estimation
3.2. 2. Homophily-Based Distance
3.3. 3. The Dynamic Social Loop
4. Experiments & Performance
4.1. Key Results:
5. Critical Insight & Conclusion