RedTweet: Bridging the Gap Between Micro-blogging and Viral Content

2015 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining

Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces RedTweet, a recommendation engine that bridges Twitter and Reddit by mapping user interests to content genres. It utilizes an ensemble of three Naive Bayesian classifiers to perform genre classification on tweets and provide personalized Reddit thread suggestions.

TL;DR

RedTweet is a recommendation engine designed to transform your Twitter activity into a curated Reddit feed. By treating recommendation as a genre classification problem, the system analyzes the "linguistic DNA" of your tweets using an ensemble of Naive Bayesian classifiers and matches them with hand-tagged subreddits. It achieves higher precision by utilizing a high-confidence threshold mechanism and oversampling techniques to balance training data.

Background & Positioning

Published during the mid-2010s social media boom (ASONAM '15), RedTweet addresses the "information overload" problem on Reddit by leveraging the "interest signals" found in Twitter's 140-character snippets. It sits at the intersection of Genre Classification and Cross-Platform Recommender Systems.

The Challenge: Content Sparsity and Imbalance

Traditional classifiers fail on tweets because:

  1. Context Scarcity: 140 characters provide very few features for a bag-of-words model.
  2. Dataset Imbalance: Classical corpora like the Brown Corpus (the gold standard for genre) are heavily skewed toward specific categories, leading to biased predictions.
  3. Ambiguity: A single piece of text often belongs to multiple genres (e.g., a "Technology" tweet might also be "News").

Methodology: The Ensemble "Bag of Decisions"

Instead of relying on a single "best" model, the authors built an ensemble. Unlike traditional voting where the majority wins, RedTweet uses a Bag Approach: every classifier's "vote" is added to a frequency tally, allowing for multi-genre representation.

The Three Pillars of the Ensemble:

  1. Classic Naive Bayesian: Trained on stemmed words and TF-IDF values from an oversampled Brown Corpus + modern Tech articles.
  2. POS-Based Classifier: Focuses on the structural components—subject phrases, object phrases, and verbs—to capture the "action" context of a tweet.
  3. Threshold Biasing Classifier: Only makes a prediction if the confidence . This acts as a stabilizer, ensuring the final profile is anchored by high-certainty predictions.

The Ensemble Architecture Fig 1: The ensemble pipeline where a tweet is processed through three distinct logic paths.

Experiments and Results

The authors validated the system by profiling high-profile Twitter users. For instance, Neil deGrasse Tyson's profile correctly peaked in "Belles Lettres" and "Learned" (Science), while Hillary Clinton's focused heavily on "News."

Key Performance Metrics:

  • Oversampling Impact: By balancing the training distribution, precision rose from 45.1% to 58.8%.
  • High Precision: The Threshold classifier reached 76.9% precision, proving that waiting for high-confidence data is better than guessing on noisy tweets.

Performance Comparison Table Table 1: Comparison of the different Naive Bayesian variations used in the study.

Critical Analysis & Conclusion

Takeaways

The "bag of decisions" is a clever way to handle multi-label classification without the complexity of traditional multi-label algorithms. It acknowledges that a user's interests are a distribution, not a single point.

Limitations

  • Manual Tagging: The 50 subreddits were hand-tagged, introducing human subjectivity (e.g., should r/creepy be 'Lore' or 'Learned'?).
  • Static Corpus: Using the 1961 Brown Corpus requires significant modern technical "patching" to stay relevant to 21st-century social media.
  • API Constraints: The dependence on limited Twitter/Reddit API calls restricts the depth of the interest profile.

Future Outlook

Today, this work could be evolved using Large Language Models (LLMs) to handle the zero-shot classification of genres, removing the need for hand-tagged subreddits and providing a more nuanced understanding of "internet slang" that Naive Bayesian models might miss.

Find Similar Papers

Try Our Examples

  • Search for recent studies that use deep learning models like BERT or RoBERTa to classify short-form social media text into the 15 Brown Corpus genres.
  • Which paper originally introduced the Brown Corpus for linguistic research, and how have modern recommender systems adapted its genre labels for digital content?
  • Explore research that applies ensemble Naive Bayesian methods to cross-domain recommendation tasks between Twitter and other multimedia platforms like YouTube or Instagram.
Contents
RedTweet: Bridging the Gap Between Micro-blogging and Viral Content
1. TL;DR
2. Background & Positioning
3. The Challenge: Content Sparsity and Imbalance
4. Methodology: The Ensemble "Bag of Decisions"
4.1. The Three Pillars of the Ensemble:
5. Experiments and Results
6. Critical Analysis & Conclusion
6.1. Takeaways
6.2. Limitations
6.3. Future Outlook