Sentiment Analysis Leveraging Emotions and Word Embeddings: A Hybrid Multilingual Approach

Expert Systems With Applications

2025-01-01
Som Gupta
Summary
Problem
Method
Results
Takeaways
Abstract

This paper proposes a hybrid vectorization methodology for multilingual sentiment analysis, combining lexicon-based emotion features with Word2Vec embeddings. Evaluated on Greek and English datasets, the approach achieves SOTA-level accuracy while significantly outperforming deep learning models like Recursive Auto-Encoders (RAE) in computational efficiency.

Executive Summary

TL;DR: This research tackles the dual challenge of language inflection complexity and computational efficiency in sentiment analysis. By merging traditional emotion-based lexicons with modern word embeddings (Word2Vec), the authors created a hybrid vectorization framework. The result is a system that achieves competitive accuracy on both English and Greek datasets while being 100x to 1000x faster than complex recursive neural architectures.

In the academic landscape, this work represents a transition from pure statistical linguistics to hybrid machine learning, proving that explicit emotional features still provide a necessary "inductive bias" that generic embeddings might miss.

The Motivation: Why Not Just Use Embeddings?

While Word2Vec and subsequent models capture semantic regularities, they often fail to grasp the nuanced emotional intensity of specific words or handle the structural complexities of high-inflection languages like Greek.

For instance, in English, a verb might have 4 forms (ask, asks, asked, asking); in Greek, a single verb can have 93 different forms. This morphological explosion leads to data sparsity. Furthermore, pure embedding models are "sentiment-blind"—they know "happy" and "sad" are related in context but don't inherently know they represent opposite polarities unless trained on massive labeled corpora.

Methodology: The Power of Hybridization

The core contribution is the Hybrid Feature Vector, constructed through two parallel tracks:

1. Enhanced Lexicon-Based Features

The authors didn't just count words. They:

  • Expanded Lexicons: Increased the Greek Sentiment Lexicon from 2,315 to 4,658 terms using thesaurus expansion.
  • Multi-Dimensional Scoring: Used EmoLex dimensions (Anger, Anticipation, Disgust, Fear, Joy, Sadness, Surprise, Trust).
  • Negation Handling: Implemented a "window" approach where words following a negation (e.g., "not") have their polarity scores reversed or doubled in the vector space.

2. Word Embedding Integration

Using Word2Vec (CBOW/Skip-gram), the system generates a dense vector representing the semantic context of the document.

3. The Concatenation

The final document vector is V_hybrid = [V_lexicon || V_word2vec]. This ensures the classifier (SVM) sees both the "what it means" (context) and "how it feels" (emotion).

Overall Framework Architecture

Experiments and SOTA Comparison

The methodology was tested against Recursive Auto-Encoders (RAE) and other traditional baselines across four datasets (English Movies, IMDB, Greek Mobile Reviews).

Key Performance metrics:

  • Greek Mobile Reviews (MOBILE-SEN): The hybrid approach reached 78.57% accuracy, outperforming pure embeddings (71.79%).
  • Efficiency: The hybrid SVM approach requires only 10^1 to 10^2 seconds per fold, whereas RAE takes 10^3 to 10^4 seconds.

Performance Comparison Table

Critical Analysis & Conclusion

The brilliance of this work lies in its pragmatism. By acknowledging that deep learning models like RAE are computationally expensive and often "black boxes," the authors offer a transparent, fast, and highly effective alternative.

Takeaway for Practitioners: If you are building a production-level sentiment engine for a non-English language or need to process "Big Data" on limited hardware, don't ignore lexicons. A hybrid approach provides the best of both worlds: the broad coverage of embeddings and the surgical precision of emotional lexicons.

Limitations: The reliance on lemmatization for Greek suggests that performance is highly dependent on the quality of the dictionary. Future work could replace Word2Vec with BERT or other Transformer models to capture even deeper contextual dependencies while retaining the emotional features.

Find Similar Papers

Try Our Examples

  • Search for recent studies that combine Transformer-based embeddings with affective lexicons for multilingual sentiment analysis.
  • Which paper first introduced the NRC Word-Emotion Association Lexicon (EmoLex), and how has it been adapted for low-resource languages in subsequent research?
  • Examine the application of hybrid vectorization techniques in real-time social media monitoring for high-inflection languages like German or Slavic languages.
Contents
Sentiment Analysis Leveraging Emotions and Word Embeddings: A Hybrid Multilingual Approach
1. Executive Summary
2. The Motivation: Why Not Just Use Embeddings?
3. Methodology: The Power of Hybridization
3.1. 1. Enhanced Lexicon-Based Features
3.2. 2. Word Embedding Integration
3.3. 3. The Concatenation
4. Experiments and SOTA Comparison
4.1. Key Performance metrics:
5. Critical Analysis & Conclusion