Hybrid Synergy: Boosting Emotion Detection with Bi-LSTM, self-Attention, and CNNs
A Comparison of Word-Embeddings in Emotion Detection from Text using BiLSTM, CNN and Self-Attention
This paper introduces a hybrid deep learning architecture combining Bi-LSTM, CNN, and Self-Attention for fine-grained emotion detection from text. The authors evaluate the model across three benchmark datasets (ISEAR, SemEval 2018/2019) and demonstrate that FastText embeddings consistently outperform Word2Vec and GloVe in this specific task.
TL;DR
Emotion detection is moving beyond simple "positive/negative" sentiment toward high-granularity psychological profiling. This paper presents a robust deep learning architecture that integrates Bi-LSTM, Self-Attention, and CNN layers to extract complex emotional features from text. By benchmarking various word embeddings, the study proves that FastText is the most effective encoder for this task, achieving significant F1-score improvements on SemEval and ISEAR datasets.
Problem & Motivation
Why is emotion detection harder than standard Sentiment Analysis? While sentiment focuses on polarity, emotions like shame, guilt, or disgust often share similar vocabularies but differ in context and intensity.
The authors identify two major gaps in prior work:
- Feature Engineering Overhead: Traditional lexicons (Keyword-based) can't handle context. A word like "killing" can mean literal murder (Fear) or "killing it at a job" (Joy).
- Architecture Limitations: RNNs are great for sequence but can lose local patterns; CNNs excel at local n-grams but ignore long-distance dependencies.
The authors' insight was to combine these into a unified pipeline where Self-Attention acts as the bridge, ensuring the model knows which parts of a sentence deserve the most "cognitive" focus.
Methodology: The Core Architecture
The proposed model follows a sophisticated stack designed to interpret text through multiple lenses:
- Embedding Layer: Converts text into dense 300D vectors. The study specifically compares Google (Word2Vec), GloVe, and FastText.
- Bi-LSTM Layer: Processes the sequence in both directions to understand the "flow" of the sentence.
- Self-Attention Mechanism: Instead of treating all words equally, it calculates weights based on token similarity, effectively highlighting the emotional "triggers" in a phrase.
- CNN & Max-Pooling: Operates on the attention-weighted matrix to extract local morphological and structural features.
- Feature Fusion: The "Local" features from the CNN are merged back with the "Sequential" features from the Bi-LSTM to provide a comprehensive vector for the final Softmax classifier.
Figure 1: The hybrid architecture showing the flow from Bi-LSTM through Self-Attention to CNN.
Experiments & Results
The model was tested on three diverse datasets: ISEAR (long-form questionnaires), SemEval 2018 (tweets), and SemEval 2019 (dialogue).
Embedding Comparison
A key contribution of this paper is the quantitative proof of FastText's superiority. Because FastText uses subword information (n-grams), it handles the slang, typos, and hashtags of social media much better than Word2Vec or GloVe.
Performance vs. Baselines
The model significantly outperformed traditional ML baselines (SVM, Naïve Bayes, Random Forest). On the SemEval 2019 Task 3, the model achieved a micro-F1 of 0.703, placing it among the top systems on the global leaderboard.
Table 1: Performance comparison on SemEval 2019 Task 3. Note the significant lead of FastText over other embeddings.
Critical Analysis & Conclusion
Takeaway
The synergy of Bi-LSTM (Global context) + Attention (Importance) + CNN (Local features) creates a highly resilient model for the noisy environment of everyday writing. The choice of embedding is not just a preprocessing step; it is a fundamental architectural decision that can sway accuracy by over 2%.
Limitations & Future Work
While the model is robust, it was published just as the "Transformer Era" (BERT/RoBERTa) began to dominate. One limitation is the static nature of the embeddings used; modern contextual embeddings might further reduce the need for the Bi-LSTM layer. The authors suggest that the next frontier is Social Robotics, where this "emotional intelligence" can help AI-human interaction feel more natural and empathetic.
Paper Source: A Comparison of Word-Embeddings in Emotion Detection from Text Code: GitHub - emofinder
