Beyond Generic Vectors: Leveraging Controversy for Robust Abusive Language Detection
Twitter-based Polarised Embeddings for Abusive Language Detection
This paper introduces a method to generate polarised word embeddings by using controversial hashtags on Twitter as proxies for polarized social communities. Using simple Linear SVM classifiers, the authors demonstrate that these embeddings are highly competitive and offer superior cross-domain portability compared to generic embeddings for abusive language detection.
TL;DR
Researchers from the University of Groningen have developed a novel way to detect toxic speech by training "polarised" word embeddings on millions of tweets filtered through controversial hashtags. While standard n-gram models often fail when moving from one platform to another, these polarised embeddings prove remarkably resilient, outperforming generic GloVe vectors in cross-domain scenarios, including technical forums like StackOverflow.
Problem & Motivation: The Portability Gap
Detecting abusive language isn't just about spotting "bad words." It involves understanding the nuances of social interaction, which vary wildly between datasets (e.g., hate speech vs. general offensiveness).
The authors identify a major Achilles' heel in current SOTA-adjacent models: domain sensitivity. A model trained on racism-related tweets often fails to detect sexism or general toxicity because it overfits to specific keywords. Traditional n-gram models are particularly prone to this "lexical bias." The authors' intuition is that by training embeddings on controversial topics—where language is naturally pushed to extremes—they can capture the "flavor" of polarising discourse that leads to abuse, regardless of the specific topic.
Methodology: Engineering Polarization
The core of this work lies in how the data for the embeddings was curated. Unlike Facebook or Reddit, Twitter doesn't have explicit "groups." To overcome this, the authors used 287 controversial keywords as proxies for communities.
The Pipeline:
- Keyword Selection: Keywords were pulled from Wikipedia’s list of controversial issues (e.g., abortion, feminism, BLM, Brexit, MAGA).
- Data Collection: 6.2 million tweets containing these hashtags were collected (132M tokens).
- Embedding Training: Two GloVe models were trained: Polarised 1 (min count 1, window 5) and Polarised 5 (min count 5, window 10).
The "Sanity Check" (Table III) reveals the effectiveness of this approach. While generic embeddings associate "immigrant" with neutral terms like "migrant," the Polarised 5 model pulls in biased terms like "illegal" or "sanctuaries," reflecting the actual linguistic landscape where abuse often occurs.

Experiments: Same Distribution vs. Cross-Domain
The authors tested their models across three major datasets: OffensEval, WH (Waseem & Hovy), and HateEval.
1. In-Domain Results
In a same-distribution scenario, n-gram models still reign supreme. This is expected as they can capture specific derogatory terms used in those specific datasets. However, the polarised embeddings remained competitive, with only a minor performance drop compared to generic GloVe.
2. Cross-Dataset & Cross-Domain (The Real Test)
The true value of polarised embeddings emerged during cross-testing. When a model trained on Twitter was asked to classify "Rude or Offensive" comments on StackOverflow, the polarised embeddings consistently beat the generic ones.

The results suggest that Polarised 5 embeddings capture a universal "adversarial" linguistic structure. Whether people are arguing about politics on Twitter or code quality on StackOverflow, the way they express toxicity shares a underlying semantic manifold that polarised embeddings successfully map.
Critical Analysis & Conclusion
Why does it work?
The success of polarised embeddings in the StackOverflow test suggests that controversy is a stylistic marker. By training on controversial hashtags, the model learns the semantics of "us vs. them" narratives, which are common precursors to abusive language.
Limitations
- Data Volume: The polarised embeddings were trained on significantly less data than generic GloVe Twitter.
- Indirect Communities: Using focal keywords is a "noisy" proxy for actual communities compared to scraping dedicated hate groups.
Takeaway
This research shifts the focus from what is being said (keywords) to the context of the interaction (polarization). For practitioners building real-world moderation systems, this highlights the necessity of using domain-aware, biased embeddings rather than "clean" generic ones to maintain performance across the shifting sands of social media discourse.
