Unmasking the Redditor: Advanced Authorship Attribution in Forum Ecosystems

Authorship Attribution using data from Reddit forum

2020-10-21
Guilherme Ramos Casimiro, Luciano Antônio Digiampietri
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates Authorship Attribution (AA) on the Reddit forum using comments from the "r/science" subreddit. It compares multiple text representations (n-grams of words, characters, and POS tags) across various classification scenarios using online learning algorithms, achieving high accuracy (often >95%) in identifying authors.

TL;DR

Researchers from the University of São Paulo analyzed over 680,000 comments from Reddit's "r/science" to determine if an author's unique "digital fingerprint" can be identified amidst thousands of users. By leveraging batch-based online learning and multi-level n-gram analysis, the study achieved near-perfect accuracy, proving that how you "decorate" your text—spacing, punctuation, and formatting—is just as revealing as the words you choose.

Context & Positioning

In the era of information warfare, Authorship Attribution (AA) has moved from historical literary analysis (like solving the mystery of The Federalist Papers) to a critical tool for fighting Fake News and Astroturfing. This paper, published at the Brazilian Symposium on Information Systems (SBSI'20), positions itself as a scalable solution for forum-style social media, where text is longer and more formatted than Twitter but more chaotic than journalism.

The Problem: The Noise of the Crowd

Most AA methods fail on social media for two reasons:

  1. Computational Complexity: Processing millions of comments using traditional "offline" models exhausts memory.
  2. Informal Syntax: Slang, abbreviations, and typos break standard linguistic models.

The authors argue that these "errors" and formatting choices (bolding, italics, tabulations) are not noise—they are features.

Methodology: Capturing the Stylistic Fingerprint

The core of the approach involves three distinct feature representations:

  • Character n-grams: Captures the use of special characters, emoticons, and white space.
  • Word n-grams: Captures vocabulary and common phrases.
  • POS Tagging: Captures the underlying grammatical structure (syntax).

The research utilized three high-scale classifiers: Stochastic Gradient Descent (SGD), Perceptron, and Passive-Aggressive. The latter was particularly effective because it uses a regularization constant () to ignore outliers, making it robust against the erratic nature of Reddit posts.

System Overview and Results Figure: ROC Curve of the Perceptron classifier for binary authorship identification.

Experiments and Key Findings

The researchers tested three scenarios:

  1. Binary: Distinguishing between the two most active users.
  2. Multiclass: Distinguishing between the top 10 most active users.
  3. One-vs-All: Identifying one specific author out of the entire subreddit.

The "Word vs. Character" Inversion

A fascinating discovery was the behavior of in different contexts:

  • Word n-grams peaked at . When increased further, the model became "too specific" (overfitting), and accuracy dropped.
  • Character/POS n-grams required or . Because characters carry less information individually, the model needed longer sequences to "see" the author’s style.

Performance Data Table: Experimental results showing Precision (P), Recall (R), and F1-Score for the top 10 authors.

Deep Insight: Why Passive-Aggressive Won

While SGD is a standard industry workhorse, it is sensitive to learning rates and hyperparameter tuning. In this study, the Passive-Aggressive classifier was the MVP. It remains "passive" when a classification is correct and becomes "aggressive" only when an error occurs, adjusting the model just enough to fix the mistake without overreacting to outliers. This makes it ideal for the "wild west" of Reddit comments.

Critical Analysis & Future Outlook

Takeaway: This work demonstrates that even in a sea of millions of users, our writing habits—down to how we use bold text or place our commas—are remarkably unique.

Limitations: The study focuses on "high-volume" authors. Identifying "casual" users who only post a few sentences remains a significant "cold-start" challenge in the AA field.

Next Steps: The authors suggest incorporating Sentiment Analysis to see if an author's emotional volatility is also a consistent identifier. This could lead to even more resilient systems for detecting botnets and coordinated disinformation campaigns.


Disclaimer: This analysis is based on the research paper "Authorship Attribution using data from Reddit forum" by Casimiro and Digiampietri (2020).

Find Similar Papers

Try Our Examples

  • Find recent studies on authorship attribution that specifically address the challenges of short-form or informal text in decentralized social networks beyond Reddit and Twitter.
  • Which paper first introduced the Passive-Aggressive algorithm for text classification, and how has its implementation for high-dimensional feature spaces evolved since SBSI'20?
  • Explore research that integrates sentiment analysis or emotional trajectory modeling with n-gram features to improve the accuracy of authorship verification.
Contents
Unmasking the Redditor: Advanced Authorship Attribution in Forum Ecosystems
1. TL;DR
2. Context & Positioning
3. The Problem: The Noise of the Crowd
4. Methodology: Capturing the Stylistic Fingerprint
5. Experiments and Key Findings
5.1. The "Word vs. Character" Inversion
6. Deep Insight: Why Passive-Aggressive Won
7. Critical Analysis & Future Outlook