HateClassify: Solving the Ambiguity of Hate Speech via Multi-Label Learning
HateClassify: A Service Framework for Hate Speech Identification on Social Media
HateClassify is a service framework designed for social media that utilizes a Sequential Convolutional Neural Network (SCNN) to identify hate speech, offensive, and non-offensive content. Its core innovation lies in treating hate speech identification as a multi-label classification problem rather than a traditional multi-class problem to better handle linguistic overlaps.
TL;DR
Hate speech detection on social media often fails because it is difficult to distinguish "hateful" content from "offensive" content using traditional mutually exclusive categories. HateClassify proposes a new CNN-based framework that shifts the paradigm to multi-label classification, boosting detection accuracy by 20% and incorporating a democratic, crowd-sourced moderation policy.
Background: The Problem of Linguistic Overlap
Major social media platforms like Twitter and Facebook face constant criticism for their "vague" rules regarding hate speech. A fundamental technical reason for this is that current AI models treat the problem as Multi-Class Classification: a tweet is either Hateful, Offensive, or Neither.
However, as the authors discover through ScatterText analysis (see below), the vocabulary overlap between "hate" and "offensive" categories can be as high as 65%. When two categories share so many core features, forcing a machine to pick just one leads to low recall and frequent misclassification.
The HateClassify Framework
The proposed framework consists of two main pillars: a Crowd-Sourced Policy and a Sequential CNN (SCNN).
1. Crowd-Sourced Policy
Instead of a centralized organization dictating what is "offensive," HateClassify allows users to vote. This allows the AI to adapt to geographical and cultural nuances, where certain terms might be considered hateful in one region but merely offensive in another.

2. High-Performance SCNN
The core of the detection engine is a Sequential Convolutional Neural Network (SCNN). This model utilizes:
- Embedding Layer: 256 dimensions to capture semantic relationships.
- Multiple Conv1D Layers: Using kernel sizes of 3, 4, and 5 to capture different n-gram patterns.
- Global Pooling & Dropout: To prevent overfitting on noisy social media text.
Why Multi-Labeling is the Game Changer
The "Eureka" moment of the paper comes from re-evaluating the task. Instead of using a Softmax layer (which forces the sum of probabilities to 1), they use sigmoid-based thresholds. This allows a piece of text to be labeled as both Hate Speech and Offensive Language.
Using the α-Evaluation metric, the authors found that this approach captures the "strictness" required for hate speech detection without being penalized by the linguistic noise of offensive slurs used in casual contexts.
Performance Evidence
As shown in the ScatterText results, the heavy clutter in segments representing both "Hate" and "Offensive" categories proves why a single-label model would naturally struggle.

In comparative experiments across three major datasets (CrowdFlower, Davidson et al., and Waseem & Hovy), the SCNN model consistently achieved higher precision than traditional SVM or Logistic Regression baselines. The shift to multi-labeling specifically addressed the "low recall" issue that plagued previous neural network attempts in this domain.
Deep Insight & Conclusion
The true value of HateClassify is its acknowledgment of semantic ambiguity. By moving away from rigid multi-class silos, the framework mirrors the complexity of human socio-linguistics.
Takeaway for Practitioners: If your NLP model is struggling with low performance on categories that "feel" similar, stop trying to separate them. Instead, adopt a multi-label architecture and let the model embrace the overlap.
Limitations: While the framework is robust, the authors note that heavily unbalanced datasets (where one class represents >85% of data) still pose a challenge for neural networks compared to traditional n-gram/SVM approaches in terms of F-measure.
