UGTE: Mastering Short-Text Emotions through the Power of User Groups

User group based emotion detection and topic discovery over short text

2019-12-12
Jiachun Feng, Yanghui Rao, Haoran Xie, Fu Lee Wang, Qing Li
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces UGTE (User Group based Topic-Emotion model), a joint framework for emotion detection and topic discovery in short texts. By integrating latent user groups and observed user characteristics (e.g., age, gender), it aggregates sparse short messages into long pseudo-documents to achieve state-of-the-art performance in topic coherence and sentiment classification.

TL;DR

Analyzing emotions in short social media posts is notoriously difficult due to "feature sparsity"—there simply aren't enough words to provide context. The UGTE (User Group based Topic Emotion) model solves this by grouping users with similar characteristics (like age or gender) and merging their posts into "pseudo-documents." This group-level insight boosts emotion detection accuracy and discovers much clearer topics than individual-level models.

Background: The Sparse Text Trap

Traditional models like Latent Dirichlet Allocation (LDA) excel at analyzing long essays but fail on short 140-character snippets. In these short texts, words rarely co-occur, making it nearly impossible for an algorithm to find a "topic." Moreover, most models ignore who is talking. A 20-year-old and a 50-year-old might use the same words but express different emotions based on their life experiences.

The Core Insight: Homophily

UGTE is built on the sociological principle of homophily: "birds of a feather flock together." By using user characteristics (Age, Sex, Country, etc.) to discover latent groups, the model can:

  1. Reduce Sparsity: Aggregate texts from a group into a long, information-rich document.
  2. Generate Portraits: Identify "Group 1" as young users feeling "joy/fear" and "Group 2" as middle-aged users feeling "sadness/guilt" about the same topic.

Methodology: The Hierarchical Approach

UGTE adds a sophisticated user-group layer to the generative process.

UGTE Model Architecture

As shown in the architecture, the model samples a Group (g) for each user, which then influences the selection of Characteristics (f), Topics (z), and Emotions (e). This structure ensures that the discovered topics are not just random word clusters, but are semantically tied to specific demographic groups.

Experimental Results: Stability and Precision

The researchers tested UGTE on the ISEAR dataset, comparing it against established baselines like MSTM and nSLTM.

1. Topic Coherence

UGTE demonstrated superior Topic Coherence (semantic meaningfulness), especially when the number of topics is small (|Z| ≤ 50). It successfully filtered out "noise words" that plague other models.

2. The Short-Text Stress Test

When forced to analyze "extremely short" texts (under 10 words), UGTE crushed the competition. While the baseline MSTM's accuracy plummeted, UGTE remained remarkably stable, proving that user metadata can "compensate" for missing words.

Performance Comparison on Short Texts

3. Case Study: Age Matters

The study found that "Age" was a critical feature. For the topic "Intimate Relationships," the model identified that younger groups associated it with "Love/Wedding" (Joy), while older groups associated it with "Injury/Death" (Sadness).

User Portraits by Age Group

Critical Analysis & Conclusion

While UGTE represents a major step forward in interpretable AI, it does have limitations. It relies on the availability of user metadata, which may be restricted due to privacy regulations (like GDPR) or platform limitations.

Future Outlook: The authors plan to merge this group-based logic with neural networks. Imagine a system that doesn't just know what was said, but predicts how different segments of society will feel about an emerging news event—all based on a few thousands tweets. That is the future UGTE is building.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize User Profiling or Demographic Embeddings to improve short text topic modeling on platforms like Twitter or Reddit.
  • Which paper first introduced the concept of "pseudo-documents" for short text aggregation, and how has this technique evolved into the User Group based approach used in UGTE?
  • How can User Group based Topic-Emotion models be extended using Graph Neural Networks (GNNs) to incorporate social relationship links alongside user characteristics?
Contents
UGTE: Mastering Short-Text Emotions through the Power of User Groups
1. TL;DR
2. Background: The Sparse Text Trap
3. The Core Insight: Homophily
4. Methodology: The Hierarchical Approach
5. Experimental Results: Stability and Precision
5.1. 1. Topic Coherence
5.2. 2. The Short-Text Stress Test
5.3. 3. Case Study: Age Matters
6. Critical Analysis & Conclusion