Decoding the Emotional Fingerprint: Enhancing Authorship Identification via Sentiment Analysis

Increasing Authorship Identification Through Emotional Analysis

2018-01-01
Ricardo Martins, José João Almeida, Pedro Rangel Henriques, Paulo Novais
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a novel framework for Authorship Identification by integrating "Emotional Profiles" into traditional writing style analysis. By combining lexicon-based sentiment extraction with Machine Learning (Naive Bayes Multinomial), the method achieves significant accuracy gains in identifying social media authors.

Executive Summary

TL;DR: Researchers from the University of Minho have demonstrated that your emotions are as unique as your vocabulary. By extracting "Emotional Profiles" from Facebook posts, they boosted the accuracy of identifying authors by over 5%, reaching an 87.41% success rate and showing massive improvements for distinct personalities.

Academic Positioning: This work bridges the gap between Sentiment Analysis and Stylometry (Authorship Attribution). While most SOTA methods focus on the "mechanics" of writing (grammar/syntax), this paper argues that the "affective" layer is a critical, underutilized dimension of an individual's writing style.

Problem & Motivation: The "Heteronym" Challenge

Authorship identification is a century-old problem, famously illustrated by the Portuguese poet Fernando Pessoa, who used multiple "heteronyms" with distinct styles to fool readers. In the digital age, influencers and politicians use social media as a "personal signature."

The authors identify a major gap in current NLP: standard tools treat text as a bag of words or a sequence of tags but ignore the emotional consistency behind the writer. For instance, while two politicians might tweet about the same climate agreement, one consistently uses "Joy" (optimism) while another skews toward "Anger" or "Negativity." This emotional bias is the "missing link" in identifying authors in short-form content.

Methodology: Building the Emotional Profile

The proposed pipeline moves beyond simple word counts. It involves a sophisticated three-track preprocessing system designed to isolate words with emotional weight.

1. The Preprocessing Pipeline

The authors use Stanford CoreNLP to perform three parallel tasks to refine the raw text:

  • POS-T (Part of Speech Tagging): Preserves only nouns, verbs, adverbs, and adjectives—the carriers of sentiment.
  • NER (Named Entity Recognition): Removes names of people and locations to prevent the model from "cheating" by identifying topics rather than styles.
  • Stopwords & Stemming: Cleans noise and normalizes word forms.

Methodology Flowchart

2. Emotion Mapping

Using the EmoLex lexicon and Plutchik’s model, the system maps words to 8 basic emotions: Anger, Anticipation, Disgust, Fear, Joy, Sadness, Surprise, and Trust.

The insight here is profound: the authors calculated the Pearson correlation between polarities and these emotions. They found that while "Joy" strongly aligns with positive polarity, "Anger" and "Fear" are the pillars of negative signatures. These distributions form the "Emotional Profile."

Experiments & Results: The "Trump" Effect

The study analyzed 2,100 posts from 8 high-profile figures, including Barack Obama, Bill Gates, and Donald Trump.

Quantifiable Gains

The baseline model (SVM on original text) achieved 82% accuracy. By introducing the emotional features and switching to a Naive Bayes Multinomial classifier, the performance jumped to 87.41%.

MetricWithout EmotionsWith Emotions
Global Accuracy82.0%87.41%
Donald Trump (Precision)31.7%76.1%
Hillary Clinton (Precision)57.1%67.6%

Results Comparison

The most striking result is the case of Donald Trump. In the baseline model, his short, unconventional writing was hard to pin down. However, his strong emotional signature (high levels of calculated "Negativity" and "Anger" relative to others) made him much easier to identify once emotional features were added.

Critical Insight & Conclusion

Takeaway

This paper proves that emotions are not noise; they are structured components of identity. In an era of AI-generated content and "Fake News," using emotional profiling could be a key defense in verifying if a message truly came from a specific human source.

Limitations & Future Work

While successful, the study relies on a lexicon-based approach, which can struggle with sarcasm or context-dependent emotional shifts. The authors suggest that the next step is measuring emotional intensity (arousal)—not just which emotion is present, but how strongly it is expressed—to further sharpen the identification "lens."

This research opens a new door for digital forensics: the way we feel is the way we write.

Find Similar Papers

Try Our Examples

  • Which recent papers explore the use of Deep Learning Transformers (like BERT or RoBERTa) specifically for emotional authorship styling compared to traditional lexicon methods?
  • Investigate the origin of Plutchik’s Wheel of Emotions in computational linguistics and how modern researchers have evolved this for short-text social media analysis.
  • Search for studies applying emotional profile identification to detect "sockpuppet" accounts or automated bots in political social media campaigns.
Contents
Decoding the Emotional Fingerprint: Enhancing Authorship Identification via Sentiment Analysis
1. Executive Summary
2. Problem & Motivation: The "Heteronym" Challenge
3. Methodology: Building the Emotional Profile
3.1. 1. The Preprocessing Pipeline
3.2. 2. Emotion Mapping
4. Experiments & Results: The "Trump" Effect
4.1. Quantifiable Gains
5. Critical Insight & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work