Cracking the Cryptic Tweet: A Stylometric Approach to Authorship Attribution
Authorship attribution for textual data on online social networks
This paper presents a comparative study of machine learning approaches for Authorship Attribution (AA) on short-form social media text (Twitter). The study evaluates multiple algorithms including Logistic Regression and SVM across varied stylometric feature sets (Lexical, Syntactic, and Word Frequency) to identify authors among a suspect pool.
Executive Summary
TL;DR: In an era of anonymous cyber-bullying and digital misinformation, identifying the real person behind a handle is vital. This research investigates how machine learning can "fingerprint" a writer's style within the 140-character limit of a tweet. By fusing n-grams and Part-of-Speech (POS) tags, the study achieves up to 75% accuracy in identifying authors from a suspect pool.
Academic Context: This work serves as a comparative benchmark in the digital forensics space, evaluating classical machine learning (SVM, Naive Bayes, Logistic Regression) against various stylometric feature combinations for short-text analysis.
The Challenge: Stylistic Sparsity
Authorship Attribution (AA) has existed for decades (e.g., identifying the Federalist Papers authors), but social media changed the rules.
- Data Sparsity: Long documents provide stable word frequencies; tweets are noisy and brief.
- Unstructured Format: Slang, emojis, and lack of standard grammar break traditional NLP tools.
- Anonymity: The ease of creating multiple accounts makes tracing unlawful activities a "needle in a haystack" problem.
Methodology: The Stylometric Pipeline
The authors treat writing style as a behavioral trait (Stylometry). Their process involves a four-stage pipeline: Data Collection → Preprocessing → Modeling → Identification.

Feature Engineering: The "Secret Sauce"
- Lexical: Character 3/4-grams and Word 1/2-grams capture morphological and thematic nuances.
- Syntactic: POS (Part-of-Speech) n-grams reveal the grammatical framework the author subconsciously uses.
- Statistical: Word Frequency Distribution helps distinguish common vs. rare word usage patterns.
To handle the resulting high-dimensional sparse matrix, the authors used PCA (Principal Component Analysis) for dimensionality reduction and Chi-Squared for selecting the -best features.
Experimental Results & Insights
The study tested 10 prolific authors with varying tweet counts (200 to 500).
Performance Winners
- Logistic Regression (LR) and Linear SVC consistently outperformed more complex kernels (like SVM-RBF) on this specific dataset scale.
- Feature Combination is Key: Using only character n-grams resulted in ~54-55% accuracy. However, combining all features (Lexical + POS + Frequency) pushed Logistic Regression to 75%.

Why Logistic Regression?
In the context of short text and limited training samples, LR acts as a strong linear baseline that avoids the overfitting often seen in Deep Learning or non-linear SVM kernels. It effectively handles the non-positive data generated during the PCA transformation.

Critical Analysis & Conclusion
Takeaways
The research proves that authorship can be identified even in micro-messages, provided that cross-domain features (syntactic + lexical) are used.
Limitations
- Topic Bias: The authors noted that accuracy can fluctuate because writing style is often influenced by the current topic (e.g., elections vs. movie reviews).
- Scalability: While 75% accuracy is impressive for 10 authors, the performance in a "wild" scenario with thousands of potential suspects remains a significant challenge (the "Open Set" problem).
Future Outlook
The next frontier in this field involves Deep Stylometry—using Large Language Models (LLMs) to capture latent stylistic embeddings that are even more robust to topic shifts and intentional "style masking."
