Cyber-Forensics: Decoding the Identity of Bullies through Computational Stylometry

15574_Computational Stylometry and Machine Learning for

Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a cyberbullying detection platform that leverages Computational Stylometry (CS) and Machine Learning (ML) to identify the gender and age of authors. Evaluated during a large-scale scientific event, the study demonstrates that stylistic patterns in offensive texts can effectively predict demographic attributes with high precision.

TL;DR

This research investigates the intersection of Computational Stylometry (CS) and Machine Learning (ML) to unmask the age and gender of individuals engaging in cyberbullying. By analyzing linguistic habits rather than just keywords, the authors developed a detection platform capable of profiling anonymous users with high precision, providing a vital tool for digital forensic investigations.

Background & Motivation

Cyberbullying is a growing social crisis, yet identifying perpetrators remains a "cat-and-mouse" game due to online anonymity. Most existing detection systems are reactive, looking for "foul words" or specific slurs. However, the authors argue that how a person writes—their unique stylistic fingerprint—is much harder to mask than what they say.

The motivation behind this work is twofold:

  1. To demonstrate that academic CS techniques can be transitioned into functional, real-time industrial applications.
  2. To prove that demographic attributes (Gender/Age) are inextricably linked to linguistic choices, especially in emotionally charged contexts like harassment.

The Core Mechanism: Computational Stylometry

At the heart of this research is the transition from simple content analysis to Authorship Profiling. The system doesn't just look for "hate speech"; it analyzes the structural DNA of the text.

Key Methodology Components:

  • Stylometric Features: These include the frequency of punctuation marks, use of capital letters (often indicating shouting or aggression), and Parts-of-Speech (POS) distribution.
  • Emotion-Labelled Graphs: Unlike traditional bag-of-words models, these features capture the syntactic structure of emotions and their specific placement within a message.
  • Machine Learning Integration: By feeding these stylometric markers into supervised ML algorithms (such as SVMs), the platform learns to differentiate, for example, the "verbal aggression patterns" of a teenage male versus an adult female.

Architecture Placeholder Note: The platform architecture integrates a web-based interface for data collection with a back-end ML engine optimized for forensic stylometry.

Experimental Insights

The research highlights critical findings from related literature and their own platform testing. For instance, while general gender detection often hovers around 64-66% accuracy in informal texts, specific forensic tasks—such as identifying predatory behavior in chat logs—can reach up to 97% accuracy when using high-level stylistic and relationship-word features.

Key Findings:

  • Age-Specific Vocabulary: Younger users tend to write more about academic subjects or homework, even when engaging in social discourse.
  • Behavioral Shifts: Work by Bogdanova et al. (referenced in the paper) shows that certain offenders follow a predictable linguistic shift from "complimentary/nice" to "emotionally unstable," which CS can detect before the abuse reaches its peak.

Experimental Results Comparison Note: Performance is measured across Precision, Recall, and F-Measure, demonstrating that stylometry significantly reduces False Positives compared to simple keyword filtering.

Critical Insight & Future Outlook

The true value of this work lies in its Inductive Bias: the assumption that language is a behavioral trait. While a bully can delete a post or avoid a "banned word," they cannot easily change their syntactic rhythm or the way they deploy punctuation under stress.

Limitations & Challenges:

  • Code-Switching: The current models struggle with users who switch languages or use heavy internet slang/emojis that don't fit standard POS tagging.
  • Data Scarcity: As bullying shifts to encrypted and ephemeral platforms (like Snapchat or Telegram), collecting large-scale corpora for training becomes more difficult.

Conclusion

By moving beyond the surface level of text, Pascucci et al. provide a blueprint for a more sophisticated generation of safety tools. The integration of Computational Stylometry into cyberbullying platforms doesn't just flag "bad" content; it provides the demographic context necessary for legal and psychological intervention.

Find Similar Papers

Try Our Examples

  • Search for recent papers published after 2020 that use Transformer-based models (like BERT or RoBERTa) for gender and age detection in cyberbullying datasets.
  • What are the primary stylometric features used in the PAN Author Profiling challenges, and how have they evolved from basic N-grams to deep learning embeddings?
  • Explore research that applies Computational Stylometry for authorship attribution in multi-modal cyberbullying tasks involving both text and image metadata.
Contents
Cyber-Forensics: Decoding the Identity of Bullies through Computational Stylometry
1. TL;DR
2. Background & Motivation
3. The Core Mechanism: Computational Stylometry
3.1. Key Methodology Components:
4. Experimental Insights
4.1. Key Findings:
5. Critical Insight & Future Outlook
5.1. Limitations & Challenges:
6. Conclusion