AD-LDA: Boosting Emotion Estimation in Tweets via Auxiliary Lexicons

Using an auxiliary dataset to improve emotion estimation in users’ opinions

2021-04-28
Siamak Abdi, Jamshid Bagherzadeh, Gholamhossein Gholami, Mir Saman Tajbakhsh
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Auxiliary Dataset-Latent Dirichlet Allocation (AD-LDA), an unsupervised topic-modeling extension designed to estimate nine distinct emotions (e.g., anger, joy, trust) in Twitter data. By incorporating a lexicon as prior knowledge during the parameter estimation process, it significantly outperforms traditional LDA in capturing the emotional nuances of short texts.

TL;DR

Researchers have developed AD-LDA (Auxiliary Dataset-Latent Dirichlet Allocation), a model that enhances emotion detection in tweets by using a specialized "auxiliary dataset" of words with known emotional weights. Unlike basic sentiment analysis (Positive/Negative), AD-LDA tracks nine specific emotions, achieving a 64% improvement in semantic coherence over standard LDA and nearly 80% accuracy in word-emotion learning.

The Motivation: Why Binary Sentiment isn't Enough

Most existing systems simplify human opinion into a boring trinity: Positive, Negative, or Neutral. However, deciding between a movie (#BlackPanther) and discussing a social movement (#MeToo) involves vastly different emotional spectra—ranging from Anticipation and Joy to Disgust and Fear.

The challenge is that Supervised Learning is expensive (requiring human-labeled tweets) and Unsupervised Learning (like traditional LDA) often gets lost in the noise of short, 280-character messages. The authors' insight was simple: Why not give the unsupervised model a "cheat sheet" of fundamental words with known emotions?

Methodology: Anchoring Local Logic in Global Knowledge

The core of the AD-LDA model is the integration of an auxiliary dataset during the Collapsed Gibbs Sampling (CGS) phase.

1. The Generative Process

The model treats every tweet as a mixture of emotions and every emotion as a mixture of words. It distinguishes between:

  • Vocabulary A: Words with a fixed, known emotion (from the auxiliary dataset).
  • Vocabulary D: New or ambiguous words whose emotions must be learned from context.

2. Model Architecture

During the learning process, when the algorithm encounters a word from the auxiliary dataset, its emotion is kept constant. This "anchor" helps the model correctly guess the emotional load of surrounding unknown words.

AD-LDA Graphical Model

Experiments: Analyzing the Pulse of Trends

The researchers tested AD-LDA on four major hashtags from 2017: #MeToo, #BlackPanther, #BoweBergdahl, and #MondayMotivation.

Key Findings:

  • Semantic Coherence: The "Coherence Score" (how well the words in a group actually relate to each other) skyrocketed by an average of 64.15%.
  • Emotion Distributions:
    • #BlackPanther was dominated by Anticipation (22.3%), naturally reflecting the hype for a future movie release.
    • #BoweBergdahl (a military court case) saw high levels of Fear (22.8%) and Disgust (17.6%).
    • #MondayMotivation predictably focused on Joy and Trust.

Comparison of Emotion Clusters

SOTA Comparison: AD-LDA vs. Conventional LDA

In the table above, we see that AD-LDA's coherence scores are much closer to zero (more positive) across all hashtags. In standard LDA, the "emotion" clusters were often a jumbled mess of unrelated terms. AD-LDA, by contrast, successfully grouped words like "hit" and "badass" under Anger, or "trailer" and "marvel" under Joy/Anticipation.

Critical Analysis & Future Outlook

Limitations: While AD-LDA is powerful, it still relies on a static auxiliary dataset. Language on social media evolves rapidly (memes, sarcasm, new slang), and a static lexicon might eventually become obsolete without dynamic updates.

The Road Ahead: The authors propose that the next step is replacing the traditional parameter estimation with Deep Learning methods and incorporating Emojis—which are essentially universal "emotion labels" already provided by users—directly into the probabilistic framework.

Takeaway

AD-LDA demonstrates that you don't always need massive labeled datasets to get high-quality results. By combining the statistical rigor of LDA with the targeted "nudge" of a small lexicon, we can gain a much more granular understanding of the public's emotional heart.

Find Similar Papers

Try Our Examples

  • Search for recent studies that integrate large-scale lexicons or knowledge graphs into Latent Dirichlet Allocation to handle short-text sparsity in social media.
  • Which paper first proposed the Joint Sentiment Topic (JST) model, and how does AD-LDA's use of an auxiliary fixed-emotion dataset specifically differ from JST's prior sentiment distribution?
  • Explore how the AD-LDA framework could be extended to include multimodal inputs, such as emojis or images, to improve emotion estimation in non-textual social media posts.
Contents
AD-LDA: Boosting Emotion Estimation in Tweets via Auxiliary Lexicons
1. TL;DR
2. The Motivation: Why Binary Sentiment isn't Enough
3. Methodology: Anchoring Local Logic in Global Knowledge
3.1. 1. The Generative Process
3.2. 2. Model Architecture
4. Experiments: Analyzing the Pulse of Trends
4.1. Key Findings:
5. SOTA Comparison: AD-LDA vs. Conventional LDA
6. Critical Analysis & Future Outlook
7. Takeaway