Beyond Follower Graphs: Clustering Social Media Users via Probabilistic Topic Modeling
Clustering Users in Micro Blogging Social Networks Using Probabilistic Topic Modeling - A Framework
This paper proposes a framework for clustering users in micro-blogging social networks like Twitter using probabilistic topic modeling. By implementing a modified Latent Dirichlet Allocation (Twitter-LDA), the methodology extracts latent interests from short-form user posts to group individuals with similar topical distributions.
TL;DR
This research introduces a methodology to categorize micro-blogging users not by who they follow (structural links), but by what they talk about (topical content). By refining Latent Dirichlet Allocation (LDA) for the constraints of Twitter, the authors provide a framework to group users based on their latent interests, overcoming the limitations of short, noisy text.
The Structural Fallacy: Why Links Aren't Enough
In platforms like Twitter, the "Following" mechanism is the standard for grouping. However, this creates a Structural Fallacy: following a journalist or a celebrity does not necessarily mean the user shares the same interests or produces similar content.
Existing methods struggle with:
- Data Sparsity: Tweets are restricted (historically 140 characters), providing very little context for standard NLP.
- Linguistic Noise: The prevalence of abbreviations, slang, and misspellings creates a "dirty" dataset that breaks traditional Latent Semantic Indexing.
- Asymmetric Relationships: Unlike Facebook's "friendship," Twitter's one-way following doesn't guarantee mutual interest.
Methodology: The Twitter-LDA Pipeline
The authors propose a robust pipeline that moves from raw RSS feeds to a "Topic Probability Table."
1. Robust Pre-processing
To tackle the noise, the framework employs:
- Stop-word Removal: Filtering 647 common English terms.
- Dictionary Normalization: Using a 45,000-word vocabulary to correct misspellings and expand abbreviations—a critical step for maintaining the integrity of the word-topic distribution.
2. Twitter-LDA: Adapted Inference
Standard LDA assumes a document is a mixture of multiple topics. In contrast, Twitter-LDA (as utilized by the authors) operates on the intuition that a single tweet is usually about one specific thing. This constraint significantly reduces the noise in the probabilistic distribution.
Note: The image above illustrates the RSS capture process, the precursor to the LDA inference.
Experiments and Results
The study focused on a specific cohort: Malaysian journalists and their followers (229 core users).
- Scale: The system extracted the maximum allowable 3,200 tweets per user via the Twitter API.
- Granularity: The LDA was configured to find 20 latent topics across the dataset.
- Outcome: The result is a User-Topic Matrix. If User A and User B both have a high probability density for "Politics" and "Technology," they are clustered together, regardless of whether they follow each other.
Critical Insight: The Value of "Topic Sensors"
The paper highlights a fascinating perspective: viewing every user as a natural event sensor. When topic modeling is applied in real-time, clusters of users shifting their focus to a specific topic (e.g., "Earthquake") can serve as a faster detection system than traditional media.
Deep Takeaway
For commercial and governmental entities, this framework moves beyond "Demographics" and into "Psychographics." It allows for:
- Market Segmentation: Micro-targeting based on actual discourse.
- Harm Prevention: Monitoring clusters discussing harmful activities.
- Platform Optimization: Helping social media providers suggest relevant content to keep users engaged.
Limitations & Future Work
While robust, the current framework is primarily focused on English. The evolution of this work would naturally involve Cross-lingual Topic Modeling to handle the multilingual nature of regions like Malaysia. Furthermore, integrating temporal analysis (how interests change over time) would add a dynamic layer to the current static clusters.
