ENRG: Bridging the Emotional Gap in Neural Response Generation
Neural Response Generation with Relevant Emotions for Short Text Conversation
The paper introduces ENRG (Emotion-Aware Neural Response Generation), a framework for Short Text Conversation (STC) that automatically determines appropriate emotions for a response and generates topically relevant, emotionally suitable content. By integrating an emotion relevance estimator with an encoder-decoder generator, the model achieves SOTA performance on the NTCIR-12 dataset.
TL;DR
While LLMs have made great strides in fluency, making chatbots "feel" human remains a challenge. This paper presents ENRG (Emotion-Aware Neural Response Generation), a framework that doesn't just generate text—it predicts the most appropriate emotional "vibe" for a response before it even starts writing. By training an emotion relevance estimator and a generator jointly, the system achieves significant gains in both response quality and emotional diversity.
Problem & Motivation: The "Cold" Bot Problem
Traditional Sequence-to-Sequence (Seq2Seq) models are trained to maximize the likelihood of the next word. This leads to safe, but "cold" and generic responses like "I don't know" or "That's interesting."
The authors identified two major gaps in prior research:
- Lack of Autonomy: Most emotional models require a user to manually select an emotion (e.g., "Respond to this post with 'Happiness'").
- Emotional Diversity: A single post (e.g., "I'm graduating today!") can legitimately trigger multiple emotions—pride, sadness (leaving friends), or surprise. Most models fail to account for this multi-modal emotional distribution.
Methodology: The ENRG Architecture
The core innovation lies in treating emotion as a latent variable that must be estimated from the context.
1. Emotion Relevance Estimator
Instead of assuming one fixed emotion, the model uses an attention-based RNN to create a probability distribution over six categories: Neutrality, Happiness, Sadness, Disgust, Surprise, and Anger. This ensures the bot understands the emotional affordance of the input.
2. Emotion-Aware Generator
The decoder is modified to accept an Emotion Embedding. This vector acts as a "style guide" for the hidden states, influencing word selection to align with the chosen sentiment.
3. Joint Learning (The Breakthrough)
The authors found that training the Estimator and Generator separately was sub-optimal. Their ENRG-Joint model shares a single encoder. This allows the model to learn hidden features that are simultaneously useful for understanding topic and sentiment.
Figure: The Joint Emotion-Aware architecture showing shared encoder parameters and the dual-optimization objective.
Experiments: Beyond Topical Relevance
The model was tested on the NTCIR-12 STC dataset using metrics like nG@1 (Normalized Gain) and P+.
Key Findings:
- Superiority of Joint Training: The
ENRG_Jointconsistently beat theENRG_Splitmodel. - Diversity Wins: Unlike many emotional models that become repetitive, ENRG Joint increased diversity scores, proving that emotional awareness helps the model explore a wider range of the vocabulary.
- Re-ranking Power: When used to re-score responses from a retrieval-based system, ENRG improved the nG@1 score by a staggering 24.7%.
Table: Comparison of ENRG variants against the baseline NRM_Loc. Note the consistent lead of the Joint model.
Critical Analysis & Conclusion
The beauty of ENRG is its explicit emotion management. By separating the "what to feel" (estimation) from "what to say" (generation) but training them together, the authors created a system that is both controllable and autonomous.
Takeaways:
- Context is Emotional: A post's meaning is not just in its words but in the emotional reaction it invites.
- Knowledge Transfer: Sharing encoder weights between emotion prediction and text generation helps the model refine its understanding of subtle social cues in short-text conversations.
Limitations: The model currently relies on a pre-trained Kim-CNN classifier to generate "ground truth" labels for training, which may introduce noise. Future work could benefit from unsupervised emotional latent space discovery or larger-scale human-labeled datasets.
