AngryBERT: Strengthening Hate Speech Detection via Emotion and Target-Aware Multitask Learning

AngryBERT: Joint Learning Target and Emotion for Hate Speech Detection

2021-01-01
Md. Rabiul Awal, Rui Cao, Roy Ka-Wei Lee, Sandra Mitrovic
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces AngryBERT, a novel multitask learning (MTL) framework that leverages BERT as a shared representation layer to improve hate speech detection. By jointly training on sentiment classification and target identification, the model achieves state-of-the-art (SOTA) results across three major benchmark datasets: DT, WZ-LS, and FOUNTA.

Executive Summary

TL;DR: AngryBERT is a multitask learning framework that enhances hate speech detection by co-training with two auxiliary tasks: Emotion Classification and Target Identification. By using a gated BERT architecture, it effectively transfers knowledge from rich auxiliary datasets to the primary task, overcoming the perennial problem of imbalanced and sparse hate speech data.

Context: This work represents a significant step in "Context-Aware" moderation, moving beyond simple keyword matching to understanding why a post is hateful (emotion) and who it attacks (target). It currently stands as a SOTA approach for BERT-based multitask text classification in the social media domain.

The "Lacking Data" Problem in Content Moderation

Detecting hate speech is notoriously difficult because hateful content is often a "needle in a haystack" within social media streams. This leads to extreme data imbalance. Furthermore, hate speech is inherently multi-faceted; it isn't just about "bad words," but about the negative sentiment (anger/disgust) and the victimization of specific groups (religion/gender).

Previous works either relied on simple data augmentation (which often introduces noise) or treated hate detection as an isolated classification task, ignoring these vital contextual signals.

Methodology: The AngryBERT Architecture

AngryBERT's core innovation lies in its Shared-Private design and its Gate Fusion mechanism.

1. The Shared Layer (Task-Invariant)

The model uses a pre-trained BERT encoder as a backbone. Because BERT is trained on massive corpora, it provides a deep linguistic understanding that is useful across all tasks (hate, emotion, and target).

2. The Private Layer (Task-Specific)

While BERT captures the general gist, each task has its own Bi-LSTM layer to extract task-specific nuances using GloVe embeddings. This ensures that the idiosyncratic "slang" of hate speech or sub-types of emotion don't get washed out by the general transformer features.

3. Gated Fusion

Instead of simple concatenation, AngryBERT uses a Gate Fusion mechanism. This acts as a learnable filter that decides, for every single word, how much "general BERT context" vs. "specific task context" should be used for the final prediction.

Model Architecture Figure 1: The AngryBERT architecture featuring the Shared BERT layer and Task-Specific Private layers.

Experiments and Results

The authors tested AngryBERT against several heavyweights, including HybridCNN, CNN-GRU, and even vanilla BERT.

  • Consistently Superior: In all three datasets (DT, WZ-LS, FOUNTA), AngryBERT secured the highest F1-scores.
  • The Power of MTL: The model outperformed "SP-MTL" and "MT-DNN," proving that its specific gated combination of emotion/target tasks is more effective than standard multitask architectures for this domain.

Experimental Results Table 1: Performance comparison across various benchmarks. Note AngryBERT's lead in Micro-F1.

Why It Works: A Look at the "Explainability"

The most compelling part of this research is how it aids explainability. In case studies, we see that for a tweet classified as "Hateful," AngryBERT simultaneously predicts:

  • Emotion: Anger, Disgust.
  • Target Group: Sexual Orientation or Religion.
  • Target Category: Individual or Group.

This provides a "reasoning chain" for moderators. If a model says a post is hateful because it targets a specific religion with "disgust," the moderation decision becomes much more transparent and easier to verify.

Critical Analysis & Conclusion

Takeaway: AngryBERT proves that hate speech detection shouldn't be a siloed task. The interplay between emotion and target-group identification provides the inductive bias necessary to handle sparse, imbalanced datasets.

Limitations: While powerful, the model relies on the availability of labeled auxiliary datasets (like SemEval). In languages or cultures where these auxiliary datasets don't exist, the model's advantage might diminish. Furthermore, the "Gate Fusion" adds computational overhead compared to a simple linear classifier on top of BERT.

Future Outlook: The next logical step is Automated Task Selection—an AI that decides which auxiliary tasks (maybe "Sarcasm" or "Irony") are most relevant for a specific dataset—and extending this into the multimodal realm where images often provide the "emotion" that the text lacks.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize Multitask Learning (MTL) specifically to address class imbalance in toxic comment classification.
  • Which paper originally proposed the "Shared-Private" architecture for MTL in NLP, and how does AngryBERT's gate fusion improve upon it?
  • Explore if the AngryBERT framework or similar BERT-based MTL models have been adapted for multimodal hate speech detection involving images and text.
Contents
AngryBERT: Strengthening Hate Speech Detection via Emotion and Target-Aware Multitask Learning
1. Executive Summary
2. The "Lacking Data" Problem in Content Moderation
3. Methodology: The AngryBERT Architecture
3.1. 1. The Shared Layer (Task-Invariant)
3.2. 2. The Private Layer (Task-Specific)
3.3. 3. Gated Fusion
4. Experiments and Results
5. Why It Works: A Look at the "Explainability"
6. Critical Analysis & Conclusion