Enhancing Topic Modeling: Leveraging User Features and Social Networks in Social Media
User Features and Social Networks for Topic Modeling in Online Social Media
This paper introduces two novel topic models, the Feature-based Topic (FT) model and the Social-based Topic (ST) model, designed to extract user interests from short, noisy social media texts. By incorporating user features and social network structures into the Latent Dirichlet Allocation (LDA) framework, the authors achieve superior performance in topic modeling and document recommendation across Epinions, Twitter, and Google+ datasets.
TL;DR
Social media text is short, noisy, and difficult to model using traditional LDA. This paper introduces the Feature Topic (FT) and Social Topic (ST) models, which treat user features and social connections as "statistical anchors" to regularize topic distributions. By assuming that "birds of a feather flock together" (homophily), these models outperform traditional LDA in both topic discovery and recommendation accuracy.
The "Data Sparsity" Dilemma
Traditional topic modeling assumes a document is long enough to provide sufficient word co-occurrence data. Social media breaks this assumption: a 140-character tweet offers very few context clues. To solve this, the authors look beyond the text. They argue that a user's location, gender, or their choice of "hashtags" (User Features), along with who they "retweet" or "plus-one" (Social Network), are powerful predictors of their latent interests.
Methodology: Regularizing the Latent Space
The authors extend the classic Latent Dirichlet Allocation (LDA) by modifying how the user-topic distribution is generated.
1. Feature Topic (FT) Model
In the FT model, a user's interest is influenced by a set of features . The weight of each feature's contribution to a specific topic is learned via a parameter . This allows the model to learn, for example, that users who use the hashtag #tech are more likely to have a distribution skewed toward "Technology" topics.
2. Social Topic (ST) Model
The ST model leverages the social graph. It assumes that a user's interest is centered around the average interest of their friends. This acts as a smoothing mechanism: if we don't have enough tweets from User A to know what they like, looking at the average distribution of their 10 closest friends provides a highly accurate "prior."
Fig 1: The graphical representation of the Feature Topic model, showing the influence of features and weights on user interests .
Experimental Insights
The research tested these models on Epinions, Twitter, and Google+.
- Perplexity: Both FT and ST showed consistently lower perplexity (better fit) than the baseline UserLDA.
- Recommendation: The models proved highly effective for document recommendation. Interestingly, the FT model was superior for "my-hits" (predicting what a user themselves would post), while the ST model excelled at "all-hits" (predicting interests within a social circle).
- Social Network Nuance: The authors found that "Follow" networks are often noisy (many people follow celebrities they don't share interests with). In contrast, "Retweet" and "Reply" networks are sparser but much better indicators of shared topic interests.
Fig 2: Precision comparison on Epinions, showcasing how the FT model (blue line) significantly improves user-specific recommendation accuracy.
Critical Analysis & Conclusion
While the FT and ST models solve the sparsity problem, they introduce a dependency on high-quality metadata. In the Google+ experiments, where features (gender/relationship status) were "weak," the performance gain was smaller. This suggests that the choice of features is crucial.
Takeaway for Practitioners: When modeling short-text data, do not rely on the text alone. Incorporating social interaction graphs as a regularizer is a robust way to overcome the lack of word co-occurrences in micro-blogs. Future work could involve Dynamic Topic Modeling to capture how these social interests evolve over time.
