EATAdaBoost: Bridging the Emotional Gap in Cross-Domain Text Analysis
A Transfer Learning Based Boosting Model for Emotion Analysis
The paper introduces EATAdaBoost (Emotion Analysis based on Transfer AdaBoost), an instance-based transfer learning framework designed to classify text emotions in domains with limited labeled data. By leveraging Word2vec and Part-of-Speech (POS) tags, it identifies "pivot instances" to bridge the source and target domains, achieving superior F1 scores compared to traditional AdaBoost and TrAdaBoost baselines.
TL;DR
Emotion analysis often fails when transitioning from one domain (e.g., blogs) to another (e.g., educational reviews) due to changing vocabulary. This paper introduces EATAdaBoost, a transfer learning model that uses semantic similarity to select "Pivot Instances"—source domain examples that use universal emotional language. This approach prevents the model from being distracted by domain-specific noise, significantly improving performance in data-scarce scenarios.
Problem & Motivation: The "Hot" Problem
Existing supervised models are rigid. A word like "hot" is positive when describing a popular actress but negative when describing a computer's cooling failure. This domain-specific polarity makes traditional machine learning models fail when they move to a new domain without new labels. While transfer learning exists, classical algorithms like TrAdaBoost often fail because they assume all mispredicted source data is noise. In reality, some source data is still useful if it relies on "common" emotional words (like "angry" or "happy").
Methodology: Identifying Pivot Instances
The core innovation lies in how the model determines which source data to trust.
- Semantic Bridge: The authors identify a set of common emotional words present in both the source and target domains.
- Cosine Similarity: Using Word2vec, they calculate the similarity between source instances and these common emotional words.
- Selective Boosting: Unlike standard AdaBoost, which increases weights for all errors, EATAdaBoost specifically increases the weight of "Pivot Instances" (high-similarity samples) when misidentified, ensuring the model masters the universal emotional cues.
In the table above, the shared emotional word "angry" allows the model to transfer knowledge across different contexts.
The Power of Part-of-Speech (POS)
The researchers found that not all words are equal. Adjectives usually carry the most emotional weight. By training a Word2vec model that includes POS tags and manually boosting the weights of adjectives, they achieved better feature representation than standard "one-hot" or TF-IDF methods.
Experiments & Results
The model was tested on two main datasets: a Chinese Blog dataset (RenCECps) and a target dataset of Educational Reviews.
Key Findings:
- Baseline Comparison: EATAdaBoost surpassed TrAdaBoost and standard AdaBoost, particularly when the number of target domain samples was very small (see the growth curves below).
- Feature Engineering: Using SVD (Singular Value Decomposition) on top of Word2vec with POS tags provided the most stable feature space.
Fig 2. demonstrates that EATAdaBoost maintains higher Mean F1 scores even when target domain data is extremely limited, effectively avoiding the "negative transfer" seen in TrAdaBoost.
Critical Analysis & Conclusion
Takeaway
EATAdaBoost proves that content-aware filtering of the source domain is superior to the "black-box" weight adjustment of earlier boosting methods. By using word embeddings as a "semantic bridge," we can mathematically define what makes a sample "transferable."
Limitations & Future Work
The authors acknowledge that while the method works well for binary or small-scale shifts, the imbalanced multi-class problem (where "Happiness" is rare compared to "None") remains a significant hurdle. Furthermore, the reliance on a predefined set of emotional words suggests that the method's success is tied to the quality of the "pivot" dictionary. Future iterations might integrate this with Large Language Models (LLMs) to automatically detect domain-invariant pivots.
Note: This research was supported by the Shanghai Key Laboratory of Trustworthy Computing.
