Decoding Human Feelings: A Semi-Supervised Approach to Multi-Label Emotion Detection in Conversations
Autoencoder for Semisupervised Multiple Emotion Detection of Conversation Transcripts
The paper introduces a semi-supervised multi-label emotion detection framework for conversation transcripts using the IMDb movie quotes corpus. It combines a custom Word2Vec-based emotion lexicon with a deep autoencoder to capture contextual nuances, achieving performance nearly comparable to human annotators.
TL;DR
Recognizing emotions in a conversation is far more complex than simple sentiment analysis. This paper presents a semi-supervised framework that uses an Autoencoder to "pre-learn" the structure of 2 million unlabeled movie quotes. By combining this with a novel Word2Vec-based emotion lexicon and multi-turn context, the system achieves a level of accuracy that rivals human annotators.
Context: Why "Happy" vs "Sad" is Not Enough
Most emotion detection research relies on Ekman’s model (six basic emotions). However, real life—and movie dialogue—is messy. We often feel "Guilt" (a mix of Joy and Fear) or "Outrage" (Surprise and Anger). Moreover, conversation is context-dependent: what you said two turns ago changes the meaning of what you say now.
Existing methods fail because:
- Data Scarcity: Human annotation for multi-label emotions is expensive and slow.
- Context Blindness: Bag-of-words models ignore the conversational flow.
- Label Dependency: They treat emotions as mutually exclusive, ignoring the correlations described in Plutchik’s theory.
Methodology: The Power of Semi-Supervision
The authors propose a five-step pipeline centered on the idea that a model should first understand how people speak before learning how they feel.
1. The Emotion Lexicon (Word2Vec + Plutchik)
Instead of a manual list of words, they used Word2Vec to map 181,276 words into a 100-dimensional space. By calculating the similarity between words and primary emotion anchors (like "Joy" or "Sadness"), they created a "fuzzy" lexicon where every word has an emotional signature.
2. Deep Autoencoder Architecture
The core of the "Semi-Supervised" magic lies in the Autoencoder.
- Step A: The model acts as an "infant," observing 2 million unlabeled sentences. It tries to compress the input into a hidden representation and reconstruct it. This forces the model to learn the "underlying structure" of language.
- Step B: The encoder weights are then used to initialize a supervised classifier, which is fine-tuned on a smaller set of 10,000 labeled utterances.
Fig 1: The architecture showing both the unsupervised reconstruction phase and the supervised classification phase.
3. Capturing Context
To solve the "Context Blindness," each utterance is represented as a 300D vector:
- [100D Current Utterance] + [100D Previous Utterance] + [100D Total Conversation Context].
Experimental Breakthroughs
The method was tested on the IMDb quotes corpus—a challenging dataset because movie dialogue often mirrors the spontaneous and complex nature of human speech.
Performance vs. Baselines
The results were striking. The Semi-Supervised Autoencoder reached an F1-score of 56.0, crushing traditional multi-label algorithms like RAkEL (36.8) and DBPNN (37.5).
Fig 2: Comparison of the Autoencoder approach against baselines and human performance.
The "Human-Level" Achievement
Human annotators achieved an F1-score of 62.6. Remarkably, the AI (at 56.0) was only slightly behind, despite the fact that humans watched the movie clips (video + audio) while the AI only read the raw text. This suggests the model successfully extracted emotional cues from text that are nearly as informative as visual/auditory signals.
Deep Insights & Future Work
The success of this work highlights two critical shifts in NLP:
- The Value of the Unlabeled: Feature learning via reconstruction (autoencoding) is an extremely powerful "warm-up" for specific downstream tasks.
- Emotional Dimensionality: Moving away from "one label per sentence" to Plutchik's multi-dimensional axes allows for a more nuanced and accurate representation of human psychology.
Limitations: The model still struggles with extremely subtle emotions like "Anticipation" and does not yet link emotions to specific entities (e.g., "I am angry at the King").
Conclusion
By treating emotion detection as a structured, multi-label problem and leveraging the vast "dark matter" of unlabeled text on the internet, this research paves the way for more empathetic AI in customer service, mental health screening, and social media analysis.
