Intensional Learning: Breaking the Manual Bottleneck in Emotion Corpus Building

Intensional Learning to Efficiently Build Up Automatically Annotated Emotion Corpora

2017-10-19
Lea Canales, Carlo Strapparava, Ester Boldrini, Patricio Martínez-Barco
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces an Intensional Learning (IL) bootstrapping framework for the automatic annotation of emotion corpora. By combining lexicon-based seed generation with compositional distributional semantics (LSA and Word2Vec), the method enables the creation of large-scale, high-quality emotion datasets without the prohibitive costs of manual labeling.

TL;DR

Building high-quality emotion datasets usually requires thousands of human hours and a high tolerance for subjective disagreement. This paper proposes a bootstrapping framework based on Intensional Learning (IL). By using a small set of "intentional" features (emotion keywords) and expanding them via Word2Vec/LSA semantic similarity, the authors can automatically generate large-scale corpora that rival manual annotations in reliability.

The Problem: The High Cost of Subjectivity

Supervised machine learning is the backbone of modern Emotion Recognition (ER). However, supervised models are only as good as their training data. In the affective domain, this creates a massive bottleneck:

  • Subjectivity: Humans often disagree on emotions due to personal background.
  • Cost & Time: Manual labeling for thousands of sentences is economically unfeasible for most researchers.
  • Domain Specificity: Current shortcuts, like "Distant Supervision" (using #hashtags), only work for social media like Twitter.

Methodology: The Two-Step Bootstrapping

The authors pivot away from "Extensional Learning" (learning from labeled examples) to Intensional Learning, where categories are defined by their properties—in this case, emotional keywords.

1. Initial Categorization (Seed Generation)

The system uses the NRC Word-Emotion Association Lexicon (Emolex) to find "seed" sentences. If a sentence contains words strongly mapped to "Joy" or "Anger," it becomes part of the initial training set. To improve coverage, they enriched Emolex with synonyms from WordNet and Oxford Thesaurus.

2. Seed Extension via Semantic Similarity

To scale the corpus, the authors used Compositional Distributional Semantic Models (CDSMs). They calculated the semantic similarity between the "seed" sentences and unlabeled data.

  • Threshold: If the similarity (measured via Cosine distance in a Word2Vec or LSA space) is > 80%, the label is propagated.
  • The Logic: If "I am feeling wonderful" is tagged as Joy, then "I am feeling marvelous" should logically share that tag.

Overall Architecture Figure 1: Overview of the initial categorization and bootstrapping workflow.

Experiments & Real-World Performance

The researchers tested their method against two famous benchmarks: the Aman Corpus (blog posts) and the Affective Text Corpus (news headlines).

Key Findings:

  • Reliability: On the Aman corpus, Cohen's Kappa (a measure of agreement) for emotions like Anger and Disgust reached over 0.90, indicating that the automatic labels are extremely consistent with human judgment.
  • Classification Accuracy: When using the automatically labeled data to train an SVM, the results were remarkably close to models trained on manually labeled data (Macro F1 of ~41% vs ~48% on Aman).
  • Lexicon Enrichment: Using Oxford synonyms significantly increased the lexicon size (from 3,462 to 10,251 words), which improved the F1-score for complex emotions like Surprise and Sadness.

Experimental Results Table Table 1: Comparison of different emotion lexicons and their attributes used in the study.

Critical Insight: Why This Works

The beauty of this approach is its genre-independence. Unlike hashtag-based supervision, this method relies on the underlying "intensional" meaning of language. By setting a high similarity threshold (80%), the authors successfully filtered the noise that usually plagues unsupervised bootstrapping.

Limitations:

  • Linguistic Nuance: The system currently ignores negation and irony (e.g., "Not happy" might be tagged as Joy because of the word "happy").
  • The 'Joy' Bias: The authors noted that "Joy" words are frequent in general language, leading to a tendency for the system to over-classify sentences as Joy.

Summary & Future Outlook

This paper provides a robust blueprint for researchers who need large affective datasets but lack the budget for a small army of annotators. By leveraging the geometric properties of word embeddings (Word2Vec) and high-quality lexicons, we can now "grow" our training sets from a handful of seeds.

The next frontier? Integrating Transformer-based embeddings (like BERT or GPT) and handling negation/irony to reach true parity with human emotional intelligence.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Large Language Models (LLMs) as zero-shot or few-shot annotators to replace traditional bootstrapping methods in emotion detection.
  • Which study first formally distinguished between Extensional Learning (EL) and Intensional Learning (IL) in the context of NLP bootstrapping, and how has that theory evolved?
  • Find research that applies the VectorSum compositional distributional semantic model to multi-modal emotion recognition tasks involving both text and audio.
Contents
Intensional Learning: Breaking the Manual Bottleneck in Emotion Corpus Building
1. TL;DR
2. The Problem: The High Cost of Subjectivity
3. Methodology: The Two-Step Bootstrapping
3.1. 1. Initial Categorization (Seed Generation)
3.2. 2. Seed Extension via Semantic Similarity
4. Experiments & Real-World Performance
4.1. Key Findings:
5. Critical Insight: Why This Works
6. Summary & Future Outlook