Boosting Social Media Topic Classification via Classifier Ensembles

Documents topic classification model in social networks using classifiers voting system

2015-10-09
Hyeoncheol Lee, Beomseok Hong, Kwangmi Ko Kim
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a topic classification system for social network documents using a Classifiers Voting System. By combining SVM, Naïve Bayes, and Decision Tree through majority voting, it achieves high accuracy in categorizing short, unstructured texts into domains like Advertising, Opinion, and Stock.

TL;DR

The explosion of social media data demands robust classification, yet the "short and messy" nature of tweets and posts cripples traditional models. This paper proposes a Classifiers Voting System that synthesizes the strengths of SVM, Naïve Bayes, and Decision Trees. Achieving an average accuracy of 93%, it provides a reliable framework for categorizing social content into Advertising, Opinions, and Financial topics.

Background: Why Short Text is a "Hard Nut to Crack"

In the era of Facebook and X (formerly Twitter), data is generated in real-time, yet it is rarely labeled. Traditional models like Latent Dirichlet Allocation (LDA) typically rely on word co-occurrence patterns that are sparse in short texts. Furthermore, social media language is "noisy"—filled with URLs, stock symbols ($Ticker), and slang.

The authors recognized that no single algorithm is a "silver bullet." For instance, while SVM excels at finding high-dimensional boundaries, Decision Trees are often better at capturing specific keyword-based logic.

Methodology: The Power of Three

The proposed system follows a rigorous pipeline: Term Extraction → Feature Generation → Ensemble Classification.

1. Feature Engineering

Instead of complex embeddings, the authors utilized a 103-dimensional vector:

  • Top 100 Frequent Words: Captured via a Boolean (exists/not exists) vector space model.
  • Special Markers: Dedicated features for $Symbols (stocks), URLs, and currency formats.

2. The Voting Mechanism

The core innovation lies in the Voting Module. Each document is processed by three separate engines:

  • Support Vector Machine (SVM): Optimized for maximum margin separation.
  • Naïve Bayes: A probabilistic approach utilizing Bayes’ theorem, effective for high-dimensionality.
  • Decision Tree: A flowchart-like logic (Iterative Dichotomiser/C4.5 style) that handles attribute-based tests.

Overall Process of the Topic Classifier

Figure 1: The architecture details how raw social data flows from extraction to the final voting decision.

Experiments and Performance

The system was tested on 2,857 human-coded documents. The results revealed an interesting phenomenon: performance is topic-dependent.

  • SVM was the champion for Advertising and Opinion topics.
  • Decision Tree outperformed others in the Stock category.

Accuracy Breakdown

By utilizing the voting system, the model effectively "hedged its bets," achieving a final accuracy between 86% and 97%.

Accuracy Results Comparison

Figure 2: Performance comparison showing that the Voting System (Final) consistently matches or exceeds the best individual classifier.

Critical Insight: Automated Labeling

One of the most valuable findings was the system's ability to self-expand. The authors tested the accuracy after automatically adding new training samples (where all three classifiers agreed). The accuracy remained stable, suggesting this method can significantly reduce the "Human-in-the-loop" cost for building large-scale social datasets.

Conclusion & Future Look

The Classifiers Voting System proves that ensemble methods are a robust defense against the volatility of social media text. While it currently requires initial human labeling, the path toward a fully automated, clustering-based system is clear. For researchers and developers in social sensing, this work underscores a fundamental truth: The consensus of multiple "weak" or specialized learners often outweighs the performance of a single "strong" model.

Limitations

  • Data Preprocessing: Still relies on human coders for the initial seed.
  • Fixed Features: Top 100 words may need frequent rotation to keep up with viral trends.

Find Similar Papers

Try Our Examples

  • Search for recent papers that improve upon majority voting systems for short text classification using Deep Learning or Transformer-based ensembles.
  • Which paper first introduced the benchmark for TF-IDF in short text analysis, and how does the 103-feature vector approach in this paper compare to modern Word2Vec or BERT embeddings?
  • Explore and find research that applies this ensemble voting methodology to multi-modal social media data, such as combined image and text classification in YouTube or TikTok descriptions.
Contents
Boosting Social Media Topic Classification via Classifier Ensembles
1. TL;DR
2. Background: Why Short Text is a "Hard Nut to Crack"
3. Methodology: The Power of Three
3.1. 1. Feature Engineering
3.2. 2. The Voting Mechanism
4. Experiments and Performance
4.1. Accuracy Breakdown
5. Critical Insight: Automated Labeling
6. Conclusion & Future Look
6.1. Limitations