HBTP: Decoding Fact vs. Fiction through the Lens of User Homogeneity

Homogeneity-Based Transmissive Process to Model True and False News in Social Networks

2019-01-30
Jooyeon Kim, Dongkwan Kim, Alice Oh
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces the Homogeneity-Based Transmissive Process (HBTP), a Bayesian nonparametric model designed to differentiate between true and false news in social networks. It leverages a novel "homogeneity index" to regulate topical similarity between users in a diffusion cascade, achieving state-of-the-art accuracy (78.1%) in news genuineness classification.

TL;DR

Researchers from KAIST have developed the Homogeneity-Based Transmissive Process (HBTP), a sophisticated Bayesian nonparametric framework that identifies fake news not just by what the news says, but by how topically similar the people sharing it are. By integrating Hierarchical Dirichlet Processes (HDP) with Gaussian Processes, the model uncovers a hidden "homogeneity index" that serves as a signature for news genuineness, outperforming existing recursive neural networks in classification tasks.

Problem & Motivation: The Echo Chamber Effect

Why does false news spread differently than the truth? Previous research suggested that "homogeneity"—the degree of shared interests among a group—is a primary driver of news dissemination. However, most SOTA models treated content (NLP) and network structure (Graph) as separate entities.

The authors' core insight is that news acts as a "bridge" for topical interest. If a news story has a high homogeneity index, it acts like a filter, propagating only through highly biased, topically aligned subgroups. If it has a low index, its spread resembles a random walk across diverse user interests. Capturing this "transmissive" property is the key to identifying misinformation.

Methodology: The Math of Influence

HBTP represents a marriage between two heavyweights of Bayesian modeling:

  1. Nonparametric Topic Modeling (HDP): Unlike standard LDA, HBTP doesn't need a pre-defined number of topics. It treats each user's interest as a probability measure that is "transmitted" to followers during a retweet.
  2. Bayesian GP-LVM: To handle the non-linear relationship between topics and the genuineness of news, the model uses a Gaussian Process to infer the latent homogeneity index ().

Model Architecture

The transmission is regulated by modifying the Gamma process construction. Specifically, the shape parameter of the Gamma distribution is scaled by the homogeneity index: This ensures that when is high, the variance between the follower's () and the predecessor's () topic distribution decreases, forcing them into topical alignment.

Model Architecture Placeholder Note: The formula above shows the transmissive Gamma process used to link user interests.

Experiments & Results: Truth is More Diverse

The model was tested on a massive Twitter dataset involving nearly 80,000 users and four categories: True, False, Non-rumor, and Unverified.

Key Findings:

  • The Signature of Truth: Unsupervised analysis revealed that "Non-rumors" have the highest homogeneity indices (shared by very specific interest groups), while "True" stories actually have the lowest homogeneity (shared by a diverse set of people). "False" news falls stubbornly in the middle.
  • Superior Accuracy: The supervised version (sHBTP) achieved an accuracy of 78.1%, significantly higher than the TD-RvNN (69.8%) and GPSTM (66.4%).

Performance Comparison Table 1: Accuracy and F1 scores across different labels. HBTP consistently leads in True and False news detection.

Topic Analysis

The model found that scientific and nature-related topics generally have low homogeneity (universal interest), whereas political and industry-related news stories exhibit high homogeneity (partisan sharing).

Topic Homogeneity Figure 1: Topics retrieved by HBTP and their corresponding latent homogeneity indices.

Critical Analysis & Conclusion

HBTP moves the field beyond simple text classification. By modeling the transmissive process of user interests, it captures the sociological reality of how information flows through social networks.

Limitations:

  • The model assumes a Directed Acyclic Graph (DAG) for retweets, which might not capture "circular" discussions or complex social loops.
  • Computational complexity is manageable (), but global inference on billion-scale networks would require further optimization of the GP-LVM inducing points.

Future Outlook: The authors suggest that combining HBTP's latent homogeneity features with the temporal-structural strengths of Recursive Neural Networks (RvNN) could yield an even more robust "Deep Hierarchical" truth-checker.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Gaussian Process Latent Variable Models (GP-LVM) specifically for social media misinformation detection or user behavior modeling.
  • Which paper first proposed the transmissive or "upstream" approach in topic modeling where multiple entities generate a single document's content, and how does HBTP iterate on that foundation?
  • Are there any studies applying homogeneity-based diffusion modeling to Cross-Modal (Image/Video) fake news detection tasks?
Contents
HBTP: Decoding Fact vs. Fiction through the Lens of User Homogeneity
1. TL;DR
2. Problem & Motivation: The Echo Chamber Effect
3. Methodology: The Math of Influence
3.1. Model Architecture
4. Experiments & Results: Truth is More Diverse
4.1. Key Findings:
4.2. Topic Analysis
5. Critical Analysis & Conclusion