Boosting Social Media Topic Classification via Classifier Ensembles
Documents topic classification model in social networks using classifiers voting system
This paper introduces a topic classification system for social network documents using a Classifiers Voting System. By combining SVM, Naïve Bayes, and Decision Tree through majority voting, it achieves high accuracy in categorizing short, unstructured texts into domains like Advertising, Opinion, and Stock.
TL;DR
The explosion of social media data demands robust classification, yet the "short and messy" nature of tweets and posts cripples traditional models. This paper proposes a Classifiers Voting System that synthesizes the strengths of SVM, Naïve Bayes, and Decision Trees. Achieving an average accuracy of 93%, it provides a reliable framework for categorizing social content into Advertising, Opinions, and Financial topics.
Background: Why Short Text is a "Hard Nut to Crack"
In the era of Facebook and X (formerly Twitter), data is generated in real-time, yet it is rarely labeled. Traditional models like Latent Dirichlet Allocation (LDA) typically rely on word co-occurrence patterns that are sparse in short texts. Furthermore, social media language is "noisy"—filled with URLs, stock symbols ($Ticker), and slang.
The authors recognized that no single algorithm is a "silver bullet." For instance, while SVM excels at finding high-dimensional boundaries, Decision Trees are often better at capturing specific keyword-based logic.
Methodology: The Power of Three
The proposed system follows a rigorous pipeline: Term Extraction → Feature Generation → Ensemble Classification.
1. Feature Engineering
Instead of complex embeddings, the authors utilized a 103-dimensional vector:
- Top 100 Frequent Words: Captured via a Boolean (exists/not exists) vector space model.
- Special Markers: Dedicated features for $Symbols (stocks), URLs, and currency formats.
2. The Voting Mechanism
The core innovation lies in the Voting Module. Each document is processed by three separate engines:
- Support Vector Machine (SVM): Optimized for maximum margin separation.
- Naïve Bayes: A probabilistic approach utilizing Bayes’ theorem, effective for high-dimensionality.
- Decision Tree: A flowchart-like logic (Iterative Dichotomiser/C4.5 style) that handles attribute-based tests.

Figure 1: The architecture details how raw social data flows from extraction to the final voting decision.
Experiments and Performance
The system was tested on 2,857 human-coded documents. The results revealed an interesting phenomenon: performance is topic-dependent.
- SVM was the champion for Advertising and Opinion topics.
- Decision Tree outperformed others in the Stock category.
Accuracy Breakdown
By utilizing the voting system, the model effectively "hedged its bets," achieving a final accuracy between 86% and 97%.

Figure 2: Performance comparison showing that the Voting System (Final) consistently matches or exceeds the best individual classifier.
Critical Insight: Automated Labeling
One of the most valuable findings was the system's ability to self-expand. The authors tested the accuracy after automatically adding new training samples (where all three classifiers agreed). The accuracy remained stable, suggesting this method can significantly reduce the "Human-in-the-loop" cost for building large-scale social datasets.
Conclusion & Future Look
The Classifiers Voting System proves that ensemble methods are a robust defense against the volatility of social media text. While it currently requires initial human labeling, the path toward a fully automated, clustering-based system is clear. For researchers and developers in social sensing, this work underscores a fundamental truth: The consensus of multiple "weak" or specialized learners often outweighs the performance of a single "strong" model.
Limitations
- Data Preprocessing: Still relies on human coders for the initial seed.
- Fixed Features: Top 100 words may need frequent rotation to keep up with viral trends.
