Decoding the Blogger’s Identity: The Power of Slang and Syntax in Stylistic Profiling

Learning Age and Gender of Blogger from Stylistic Variation

2009-01-01
Mayur Rustagi, R. Rajendra Prasath, Sumit Goswami, Sudeshna Sarkar
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a stylistic analysis framework for predicting the age and gender of bloggers using a Naive Bayes classifier. It introduces the use of "slang words" (non-dictionary terms) and average sentence length as core features, achieving a remarkable accuracy of up to 97.41% for gender and 95% for age group detection.

TL;DR

Can the way you misspell words or the length of your sentences reveal your age and gender? This paper demonstrates that it can. By analyzing "non-dictionary" words—slangs, chat abbreviations, and intentional typos—alongside sentence structures, researchers achieved a staggering 97.41% accuracy in gender detection and 95% in age group classification, proving that our subconscious writing habits are digital fingerprints.

Background & Motivation: Beyond the Dictionary

In the world of Information Retrieval (IR), knowing who is writing is as important as what is being written. Traditional linguistic analysis often focuses on formal grammar. However, the blogosphere is a "Wild West" of language—informal, unedited, and rife with deviations from the norm.

The authors argue that "Style" is a subconscious habit. While we choose our topics (Content), our choice of specific words or sentence lengths (Style) happens under the radar. The core insight here is that slang is not noise; it is data. A teenager’s use of "soooo" or "bored" carries demographic signals that "standard" linguistic tools often discard.

Methodology: The Anatomy of Digital Style

The study utilized a massive corpus of 92,381 blog files, primarily from MySpace, categorized into 10s, 20s, 30s, and 40s age groups. The methodology rests on two innovative pillars:

1. The Slang Index (Non-Dictionary Words)

The researchers used Ispell to filter out every word not found in a standard dictionary. This left them with a "treasure trove" of slangs, smiley faces, and abbreviations. They filtered these for high-frequency usage (occurrence > 50) where the usage between genders was at least double. This resulted in a specialized 52-word feature list.

2. Syntactic Variation (Sentence Length)

They tracked the average sentence length across demographics. While formal literature traditionally sees sentence length increase with education and age, the blogosphere presents a more chaotic trend.

Model Feature Analysis Figure 1: Comparison of average sentence length and slang usage across age groups.

Experiments & Results: A Remarkable Leap

The results confirm a clear stylistic divide:

  • Teenagers (10s): Significantly higher usage of out-of-dictionary words and generally shorter, more emotive sentence structures.
  • Adults (30s): Lower slang usage and distinct professional/lifestyle keywords (e.g., "tax," "campaign," "systems").

Performance Gains

When the authors augmented traditional content words with their new slang markers, the accuracy of the Naive Bayes classifier skyrocketed:

  • Gender Accuracy: Jumped from a baseline of ~80% to 97.41%.
  • Age Accuracy: Achieved 95% when distinguishing between those with "remarkable age differences" (10s vs. 30s).

Confusion Matrix Table: The confusion matrix highlighting the high precision in male vs. female classification.

Critical Insight: Why Sentence Length is Tricky

Interestingly, the authors found that while average sentence length is a "remarkable feature," it isn't a silver bullet. Because blogs are informal, the variation within a single age group is vast. They suggest that true trends in sentence length can only be captured by tracking a single user over decades—a "longitudinal" study—to see how their personal style evolves as they age.

Conclusion & Future Outlook

This paper elevates slang from "poor grammar" to a "demographic marker." The implications are broad:

  • Marketing: Highly targeted ads based on the vibe of a blog rather than just keywords.
  • Forensics: Identifying the persona behind anonymous digital content.
  • Linguistics: Tracking the "evolution and death" of slangs over time.

The authors’ next frontier? Using these same features to predict geographical location and ethnic groups, as slang is often as much about where you are as who you are.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize Deep Learning or Transformers to classify author demographics based on informal linguistic slangs or emojis.
  • Which seminal paper first defined "stylistic variation" in computational linguistics, and how does this paper's focus on non-dictionary words evolve those original theories?
  • How has modern LLM-based stylometry been applied to detect geographical or ethnic backgrounds in social media datasets compared to the Naive Bayes approach used here?
Contents
Decoding the Blogger’s Identity: The Power of Slang and Syntax in Stylistic Profiling
1. TL;DR
2. Background & Motivation: Beyond the Dictionary
3. Methodology: The Anatomy of Digital Style
3.1. 1. The Slang Index (Non-Dictionary Words)
3.2. 2. Syntactic Variation (Sentence Length)
4. Experiments & Results: A Remarkable Leap
4.1. Performance Gains
5. Critical Insight: Why Sentence Length is Tricky
6. Conclusion & Future Outlook