Beyond Keywords: Bridging Psychology and AI to Solve Online Conflictual Language
Automatic Identification of Harmful, Aggressive, Abusive, and Offensive Language on the Web: A Survey of Technical Biases Informed by Psychology Literature
This paper presents a comprehensive survey of Online Conflictual Language (OCL) detection, proposing a reconciled taxonomy informed by psychology literature. It systematically maps technical gaps in current NLP systems and provides a framework for addressing biases in dataset engineering and machine learning models.
TL;DR
The automatic detection of hate speech and cyberbullying is often treated as a solved problem in "leaderboard" culture, but it remains broken in the real world. This survey by Balayn et al. highlights a massive terminological and conceptual mismatch between how computer scientists build models and how humans experience conflict. By introducing a psychology-informed taxonomy, the authors expose why current SOTA models fail to generalize and offer a roadmap for building more robust, fair, and context-aware moderation systems.
The "Decontextualization" Crisis
Most current NLP models for abusive language operate in a vacuum. A model sees a string of text, checks for "toxic" tokens, and spits out a probability.
Why is this a problem? Psychology tells us that the perception of harm is inherently subjective. A phrase used jokingly between friends (context) is vastly different from the same phrase used as a slur by a stranger. Current technical pipelines ignore this by:
- Uniforming subjective labels: Using majority voting to crush dissenting annotator opinions, which often silences the perspectives of marginalized groups.
- Ignoring the Observer: Failing to account for the fact that a woman might perceive a tweet as more aggressive than a man would, based on lived experience.
Methodology: A Multi-Dimensional Taxonomy
The core contribution of this work is a reconciled taxonomy that moves beyond "Hate Speech" as a catch-all term. The authors define Online Conflictual Language (OCL) through seven distinct anchors.

This framework allows us to distinguish between:
- Aggression: Driven by parental/behavioral intent to harm.
- Offensive Language: Centered on the target's characteristics.
- Abusive Language: Defined by the style (e.g., profanity) rather than the target.
Technical Biases in the Pipeline
The paper systematically deconstructs the machine learning pipeline to find where "biases" creep in:
1. Data Collection (The Retrieval Bias)
Most datasets are built using keyword-based sampling. If you only collect data containing "bad words," your model becomes a glorified profanity filter. It will fail to detect "coded" hate speech or microaggressions that use polite but exclusionary language.
2. Annotation (The "Majority Rule" Bias)
Current practices favor Fleiss’ Kappa or high agreement metrics. However, for OCL, disagreement is often data, not noise. By forcing a single ground truth, we introduce Aggregation Bias, where the "average" (often majority-group) opinion becomes the model's objective reality.
3. Feature Engineering (The Context Mismatch)
While deep learning (CNNs/RNNs) has improved accuracy, most models still only look at the textual content. As shown in the study, only a fraction of papers utilize metadata like user history, network centrality, or conversation threading.

Critical Insight: The "Generalization" Wall
One of the most sobering results discussed is the performance drop when models move from "Laboratory" datasets to "Deployment" data. A model achieving a 70 F1-score on its native test set can drop to a mere 21.1 F1 when tested on a different platform. This is a direct result of the technical biases listed above—the model learns the specific quirks of the dataset rather than the general nature of conflict.
Future Outlook: Building Truly Mature Systems
To move forward, the authors advocate for:
- Human-in-the-Loop 2.0: Using "CrowdTruth" approaches that embrace disagreement.
- Counter-speech Generation: Moving from simple "delete" moderation to more sophisticated linguistic interventions.
- Intersectional Evaluation: Ensuring models don't just work "on average" but are fair across different demographic slices (Gender, Ethnicity, Age).
Conclusion
Balayn et al. remind us that OCL detection is not just a "math problem" to be optimized. It is a social science challenge that requires technical humility. We cannot build safer online spaces until our models understand the Why and the Who, not just the What.
