Cracking the Cryptic Tweet: A Stylometric Approach to Authorship Attribution

Authorship attribution for textual data on online social networks

2017-08-01
Ritu Banga, Pulkit Mehndiratta
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a comparative study of machine learning approaches for Authorship Attribution (AA) on short-form social media text (Twitter). The study evaluates multiple algorithms including Logistic Regression and SVM across varied stylometric feature sets (Lexical, Syntactic, and Word Frequency) to identify authors among a suspect pool.

Executive Summary

TL;DR: In an era of anonymous cyber-bullying and digital misinformation, identifying the real person behind a handle is vital. This research investigates how machine learning can "fingerprint" a writer's style within the 140-character limit of a tweet. By fusing n-grams and Part-of-Speech (POS) tags, the study achieves up to 75% accuracy in identifying authors from a suspect pool.

Academic Context: This work serves as a comparative benchmark in the digital forensics space, evaluating classical machine learning (SVM, Naive Bayes, Logistic Regression) against various stylometric feature combinations for short-text analysis.

The Challenge: Stylistic Sparsity

Authorship Attribution (AA) has existed for decades (e.g., identifying the Federalist Papers authors), but social media changed the rules.

  • Data Sparsity: Long documents provide stable word frequencies; tweets are noisy and brief.
  • Unstructured Format: Slang, emojis, and lack of standard grammar break traditional NLP tools.
  • Anonymity: The ease of creating multiple accounts makes tracing unlawful activities a "needle in a haystack" problem.

Methodology: The Stylometric Pipeline

The authors treat writing style as a behavioral trait (Stylometry). Their process involves a four-stage pipeline: Data Collection → Preprocessing → Modeling → Identification.

Process of Authorship Analysis

Feature Engineering: The "Secret Sauce"

  1. Lexical: Character 3/4-grams and Word 1/2-grams capture morphological and thematic nuances.
  2. Syntactic: POS (Part-of-Speech) n-grams reveal the grammatical framework the author subconsciously uses.
  3. Statistical: Word Frequency Distribution helps distinguish common vs. rare word usage patterns.

To handle the resulting high-dimensional sparse matrix, the authors used PCA (Principal Component Analysis) for dimensionality reduction and Chi-Squared for selecting the -best features.

Experimental Results & Insights

The study tested 10 prolific authors with varying tweet counts (200 to 500).

Performance Winners

  • Logistic Regression (LR) and Linear SVC consistently outperformed more complex kernels (like SVM-RBF) on this specific dataset scale.
  • Feature Combination is Key: Using only character n-grams resulted in ~54-55% accuracy. However, combining all features (Lexical + POS + Frequency) pushed Logistic Regression to 75%.

Experimental Results Comparison

Why Logistic Regression?

In the context of short text and limited training samples, LR acts as a strong linear baseline that avoids the overfitting often seen in Deep Learning or non-linear SVM kernels. It effectively handles the non-positive data generated during the PCA transformation.

Analysis of accuracy with Logistic Regression

Critical Analysis & Conclusion

Takeaways

The research proves that authorship can be identified even in micro-messages, provided that cross-domain features (syntactic + lexical) are used.

Limitations

  • Topic Bias: The authors noted that accuracy can fluctuate because writing style is often influenced by the current topic (e.g., elections vs. movie reviews).
  • Scalability: While 75% accuracy is impressive for 10 authors, the performance in a "wild" scenario with thousands of potential suspects remains a significant challenge (the "Open Set" problem).

Future Outlook

The next frontier in this field involves Deep Stylometry—using Large Language Models (LLMs) to capture latent stylistic embeddings that are even more robust to topic shifts and intentional "style masking."

Find Similar Papers

Try Our Examples

  • Find recent papers from 2023-2025 that apply Transformer-based models like BERT or RoBERTa to the problem of short-text Authorship Attribution on Twitter.
  • Which seminal paper first introduced the "Source Code Author Profile" (SCAP) method, and how has its approach to character n-grams influenced modern stylometry for social media?
  • Explore research that applies multi-modal authorship attribution, combining textual stylometry with user metadata (timestamps, interaction patterns) in digital forensics.
Contents
Cracking the Cryptic Tweet: A Stylometric Approach to Authorship Attribution
1. Executive Summary
2. The Challenge: Stylistic Sparsity
3. Methodology: The Stylometric Pipeline
3.1. Feature Engineering: The "Secret Sauce"
4. Experimental Results & Insights
4.1. Performance Winners
4.2. Why Logistic Regression?
5. Critical Analysis & Conclusion
5.1. Takeaways
5.2. Limitations
5.3. Future Outlook