Selective Listening: Advancing Informative Tweet Detection via Weighted Naive Bayes

Informative vs. Non-informative Short Message Detection in Social Networks

2017-08-01
Konstantinos Giannakopoulos
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a Weighted Binary Multinomial Naive Bayes (BMNB) variation designed to classify tweets into "informative" (public interest/news) and "non-informative" (personal/noise) categories. By incorporating specialized feature weighting for hashtags and user mentions alongside a data-driven prior distribution, the method significantly enhances the retrieval of useful social media content.

TL;DR

In the deluge of social media, separating the signal (news, trends, public events) from the noise (personal chatter, mood updates) is a critical preprocessing step. This paper proposes a Weighted Binary Multinomial Naive Bayes (BMNB) model that treats hashtags and user mentions as "premium" features. By replacing generic smoothing with a data-driven prior, the model achieves a recall of up to 94%, ensuring that almost no informative content is lost while effectively downsizing massive datasets.

Problem & Motivation: The Sparsity Trap

Why is Twitter data so hard to clean? The authors identify three primary "sparsity traps":

  1. Short Length: Traditional Topic Models like LDA fail because there isn't enough word co-occurrence in 140-280 characters.
  2. Frequency Paradox: In short messages, TF-IDF weights are often meaningless because terms rarely repeat within a single document.
  3. Linguistic Chaos: Slang, typos, and multilingual tokens make standardized dictionary indexing highly inefficient.

The core motivation here isn't just classification—it's data reduction. If we can accurately discard "personal noise" (non-informative tweets) without losing public updates, we drastically reduce dimensionality for downstream tasks like recommendation engines or trend detection.

Methodology: Beyond Uniform Smoothing

The authors' "Secret Sauce" involves two key surgical modifications to the standard BMNB model:

1. The Physics of Weighting

Instead of treating every token equally, the model categorizes tokens into three buckets: Words, Hashtags, and Mentioned Users. Through empirical testing, they discovered that user mentions are the strongest indicators of informativeness (often pointing to journalists, celebrities, or organizations). The conditional probability is modified as: Where is the weight assigned to the specific feature type.

2. Prior Distribution vs. Laplace

Standard Naive Bayes uses Laplace (+1) smoothing, which assumes a Uniform Prior. The authors argue this is suboptimal. Instead, they sample 10% of the training data to estimate a Dirichlet Prior that reflects the actual distribution of terms in the specific corpus, making the model much more "aware" of the environment it is operating in.

Model Architecture/Formula

Experiments & Results: High Recall, Low Noise

The authors tested their approach on two independent datasets (Dataset A: ~20k tweets; Dataset B: ~10k tweets).

Key Findings:

  • The Weight of a User: The best performance was achieved when mentioned users were weighted significantly higher (e.g., ) than standard words ().
  • Recall is King: Since the goal is filtering, the authors prioritized scores (weighting recall over precision). Their model maintained a recall of ~92-94%, meaning it barely missed any "useful" tweets.
  • Efficiency: The model successfully identified and labeled over 6,000 tweets in Dataset A as non-informative. For a production system, this means a 30% reduction in data volume with negligible information loss.

Performance Comparison Fig. 1. TP comparison: The weighted model (green) consistently captures more informative messages than the baseline BMNB.

Critical Analysis & Conclusion

The value of this work lies in its simplicity and interpretability. While modern LLMs could solve this classification task, a Weighted BMNB is computationally "cheap" and can run in real-time on massive firehose streams without GPU clusters.

Limitations:

  • Manual Weighting: The weights were determined manually. Future iterations could benefit from automated hyperparameter optimization or attention-based weighting.
  • Context Blindness: BMNB still treats words as a "bag," potentially missing nuanced sarcasm or context-dependent informativeness.

Takeaway: This paper proves that even "old" algorithms like Naive Bayes can be highly competitive if we inject domain-specific structural knowledge (like the inherent value of a Hashtag) into the mathematical prior.

Find Similar Papers

Try Our Examples

  • Search for recent papers that integrate social graph metadata with transformer-based models for short text classification on Twitter or X.
  • Which study first introduced the Binary Multinomial Naive Bayes model for text categorization, and how has the "weighted feature" approach evolved since then?
  • Explore how the distinction between informative and non-informative content is used in modern real-time event detection systems for social media.
Contents
Selective Listening: Advancing Informative Tweet Detection via Weighted Naive Bayes
1. TL;DR
2. Problem & Motivation: The Sparsity Trap
3. Methodology: Beyond Uniform Smoothing
3.1. 1. The Physics of Weighting
3.2. 2. Prior Distribution vs. Laplace
4. Experiments & Results: High Recall, Low Noise
5. Critical Analysis & Conclusion