Hybrid Synergy: Boosting Emotion Detection with Bi-LSTM, self-Attention, and CNNs

A Comparison of Word-Embeddings in Emotion Detection from Text using BiLSTM, CNN and Self-Attention

2019-06-06
Marco Polignano, Pierpaolo Basile, Marco de Gemmis, Giovanni Semeraro
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a hybrid deep learning architecture combining Bi-LSTM, CNN, and Self-Attention for fine-grained emotion detection from text. The authors evaluate the model across three benchmark datasets (ISEAR, SemEval 2018/2019) and demonstrate that FastText embeddings consistently outperform Word2Vec and GloVe in this specific task.

TL;DR

Emotion detection is moving beyond simple "positive/negative" sentiment toward high-granularity psychological profiling. This paper presents a robust deep learning architecture that integrates Bi-LSTM, Self-Attention, and CNN layers to extract complex emotional features from text. By benchmarking various word embeddings, the study proves that FastText is the most effective encoder for this task, achieving significant F1-score improvements on SemEval and ISEAR datasets.

Problem & Motivation

Why is emotion detection harder than standard Sentiment Analysis? While sentiment focuses on polarity, emotions like shame, guilt, or disgust often share similar vocabularies but differ in context and intensity.

The authors identify two major gaps in prior work:

  1. Feature Engineering Overhead: Traditional lexicons (Keyword-based) can't handle context. A word like "killing" can mean literal murder (Fear) or "killing it at a job" (Joy).
  2. Architecture Limitations: RNNs are great for sequence but can lose local patterns; CNNs excel at local n-grams but ignore long-distance dependencies.

The authors' insight was to combine these into a unified pipeline where Self-Attention acts as the bridge, ensuring the model knows which parts of a sentence deserve the most "cognitive" focus.

Methodology: The Core Architecture

The proposed model follows a sophisticated stack designed to interpret text through multiple lenses:

  1. Embedding Layer: Converts text into dense 300D vectors. The study specifically compares Google (Word2Vec), GloVe, and FastText.
  2. Bi-LSTM Layer: Processes the sequence in both directions to understand the "flow" of the sentence.
  3. Self-Attention Mechanism: Instead of treating all words equally, it calculates weights based on token similarity, effectively highlighting the emotional "triggers" in a phrase.
  4. CNN & Max-Pooling: Operates on the attention-weighted matrix to extract local morphological and structural features.
  5. Feature Fusion: The "Local" features from the CNN are merged back with the "Sequential" features from the Bi-LSTM to provide a comprehensive vector for the final Softmax classifier.

Model Architecture Figure 1: The hybrid architecture showing the flow from Bi-LSTM through Self-Attention to CNN.

Experiments & Results

The model was tested on three diverse datasets: ISEAR (long-form questionnaires), SemEval 2018 (tweets), and SemEval 2019 (dialogue).

Embedding Comparison

A key contribution of this paper is the quantitative proof of FastText's superiority. Because FastText uses subword information (n-grams), it handles the slang, typos, and hashtags of social media much better than Word2Vec or GloVe.

Performance vs. Baselines

The model significantly outperformed traditional ML baselines (SVM, Naïve Bayes, Random Forest). On the SemEval 2019 Task 3, the model achieved a micro-F1 of 0.703, placing it among the top systems on the global leaderboard.

Experimental Results Table 1: Performance comparison on SemEval 2019 Task 3. Note the significant lead of FastText over other embeddings.

Critical Analysis & Conclusion

Takeaway

The synergy of Bi-LSTM (Global context) + Attention (Importance) + CNN (Local features) creates a highly resilient model for the noisy environment of everyday writing. The choice of embedding is not just a preprocessing step; it is a fundamental architectural decision that can sway accuracy by over 2%.

Limitations & Future Work

While the model is robust, it was published just as the "Transformer Era" (BERT/RoBERTa) began to dominate. One limitation is the static nature of the embeddings used; modern contextual embeddings might further reduce the need for the Bi-LSTM layer. The authors suggest that the next frontier is Social Robotics, where this "emotional intelligence" can help AI-human interaction feel more natural and empathetic.


Paper Source: A Comparison of Word-Embeddings in Emotion Detection from Text Code: GitHub - emofinder

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Transformer-based models like BERT or RoBERTa for emotion detection to compare against this Bi-LSTM/CNN hybrid approach.
  • Which paper first introduced the FastText embedding method, and what specific characteristics of character n-grams make it particularly effective for informal or misspelled text?
  • Explore research that applies self-attention and CNN hybrid architectures to multimodal emotion detection, specifically combining text with facial expressions or audio features.
Contents
Hybrid Synergy: Boosting Emotion Detection with Bi-LSTM, self-Attention, and CNNs
1. TL;DR
2. Problem & Motivation
3. Methodology: The Core Architecture
4. Experiments & Results
4.1. Embedding Comparison
4.2. Performance vs. Baselines
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work