Beyond Keywords: A Heuristic-Lexicon Hybrid for Decoding Digital Emotions
Lexicon and Heuristics Based Approach for Identification of Emotion in Text
This paper introduces a hybrid emotion recognition framework combining WordNet-derived lexicons with a heuristic rule engine. It classifies informal text into Ekman's six basic emotion categories (Happiness, Sadness, Anger, Fear, Disgust, Surprise) and achieves a competitive average F-measure of 0.83 on Twitter data.
TL;DR
Recognizing human emotion in the wild—specifically on social media—requires more than just a dictionary. This paper presents a hybrid approach that marries WordNet-based lexicons with informal emoticons and a heuristic engine to handle negations and intensity. By focusing on the structural "vibe" of a tweet (casing, punctuation, and slangs), the model reaches an impressive 0.83 F-measure without requiring the massive overhead of deep learning.
Problem & Motivation: The Chaos of Social Text
Why is emotion detection so hard? In a formal academic paper, a "keyword spotting" approach might work. But on Twitter, users express themselves through:
- Punctuation: "!!!" adds intensity.
- Visual Cues: Emoticons like
:)orROFLcarry more weight than words. - Negations: A single "not" can invert a sentence's entire emotional vector.
- Slangs: Terms like "OMG" or "yuck" are often missing from standard linguistic databases.
Current SOTA methods often rely on heavy Machine Learning (SVM, Naive Bayes), which necessitates massive annotated datasets and struggle with the contextual "flip" of negations.
Methodology: The Three-Pillar Approach
The authors propose a system that operates on three distinct levels to ensure no contextual nuance is lost.
1. The Iterative Lexicon (WordNet)
Instead of manually labeling thousands of words, the authors started with small sets of seed words (e.g., "satisfaction" for Happiness). Using Algorithm 1, they crawled WordNet synsets for three iterations. Each step away from the seed word "penalized" the emotion weight by 10%, resulting in a nuanced dictionary of over 3,700 weighted terms.
2. The Emoticon & Slang Layer
Since WordNet doesn't speak "Internet," the authors manually curated a dataset of popular emoticons and abbreviations (OMG, LOL) from Skype and Facebook, assigning them direct emotional vectors.
3. The Heuristic Engine (The Secret Sauce)
This is where the model moves from static matching to dynamic understanding. The logic handles:
- Intensity: Uppercase words and exclamation marks increase weights by 50% and 20%, respectively.
- The Switch: If a negation (no, not) is detected, the weights for positive and negative emotions are swapped.
- Adverb Modifiers: Words like "extremely" act as multipliers for the following keyword.
Note: Table II shows how different words like "frustrated" or "sudden" are mapped to specific emotion vectors.
Experiments & Results
The system was tested on 150 manually annotated tweets. While 73% accuracy might seem modest compared to modern LLMs, the Precision and Recall reveal a much higher level of reliability in specific categories.
Key Takeaway: The model is exceptionally good at identifying 'Happiness' (Precision 0.95) and 'Disgust' (Recall 0.92).
The high F-measure across the board (avg 0.83) suggests that the heuristic rules for negations and intensity modifiers effectively bridge the gap where simple keyword spotting usually fails.
Critical Insight: The Limitation of the "Flip"
One significant takeaway from the authors' self-critique is the complexity of nuanced negations. Currently, the model "flips" the emotion. However, a phrase like "not very disappointed" doesn't necessarily mean "Happy"; it might just mean "slightly Sad." Future iterations of this work would need to move toward a more gradient-based weight adjustment for negations rather than a binary switch.
Conclusion
This paper serves as a reminder that before jumping into multi-billion parameter models, understanding the heuristics of language—how we use capitalization, punctuation, and social symbols—can provide a robust, transparent, and computationally efficient baseline for emotion recognition.
