Beyond Words: Leveraging Social Milieu for Arabic Hate Speech Detection

Beyond Hostile Linguistic Cues: The Gravity of Online Milieu for Hate Speech Detection in Arabic

2019-09-12
Arijit Ghosh Chowdhury, Aniket Didolkar, Ramit Sawhney, Rajiv Ratn Shah, R. Shah
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a novel framework for religious hate speech detection in Arabic tweets by integrating social context with linguistic features. Using Node2Vec for network embeddings and an LSTM-CNN architecture for text, the method achieves a state-of-the-art F1-score of 0.78 on the Albadi dataset.

TL;DR

This research moves beyond simple keyword matching and text analysis for Arabic hate speech detection. By integrating Node2Vec user embeddings (from follower and retweet graphs) with an LSTM-CNN text classifier, the authors demonstrate that who you interact with is just as informative as what you say. They achieved a new SOTA on religious hate speech detection in the Arabic Twittersphere.

The Problem: The "Hostile Term" Trap

Traditional NLP systems for hate speech detection often fall into a trap: they become over-sensitive to specific "hostile linguistic cues." In Arabic, a language with complex morphology and diverse dialects, this leads to low precision.

The authors argue that hate speech doesn't exist in a vacuum. It is a social phenomenon driven by Homophily—the tendency of individuals to associate with others who share similar views. Prior works ignored this "online milieu," missing out on the rich context provided by a user's social network.

Methodology: Fusing Language and Social Graphs

The proposed framework is a two-stream architecture that fuses textual and social features.

1. Dual-Track Word Embeddings

To capture the nuances of Arabic, the authors use massive 600-dimensional embeddings created by concatenating:

  • AraVec: Pre-trained on a large Arabic Twitter corpus.
  • Domain-Specific Embeddings: Trained specifically on the religious hate speech dataset to capture task-specific vocabulary.

2. Social Network Embeddings (The "Milieu")

The core innovation lies in the use of Node2Vec. The authors construct two undirected graphs:

  • Follower Graph: Capturing long-term interest and social circles.
  • Retweet Graph: Capturing active propagation of ideas and immediate influence. These graphs are projected into a 64-dimensional latent space where users with similar interaction patterns are placed closer together.

Model Architecture Figure 1: The proposed architecture concatenating LSTM-CNN text features with Node2Vec user embeddings.

Experimental Results & Insights

The researchers tested various deep learning backbones (GRU, LSTM, CNN). The results consistently showed that adding network information improves every single baseline.

ArchitectureAccuracyF1-ScoreAUROC
Albadi-GRU (Text Only)0.770.750.84
Albadi-GRU + Network0.790.750.85
LSTM + CNN + Network0.790.780.86

Performance Comparison Table 1: Performance metrics showing the superiority of the combined approach.

Key Observations:

  • Complementary Strengths: Text features capture the intent of a specific tweet, while network features capture the character of the user.
  • Interconnected Circles: The success of the network embeddings confirms that users who proliferate hate speech tend to form tightly-knit, interconnected social clusters.

Final Thoughts & Future Directions

The "Gravity of Online Milieu" is a powerful signal. While this 2019 work used Node2Vec, the logical next step in today's landscape would be the application of Graph Convolutional Networks (GCNs) or Graph Transformers to model these relationships even more dynamically.

The main limitation remains the volatility of Twitter data (now X), as the authors noted they could only retrieve ~70% of the original dataset due to tweet deletions. This suggests that future systems need to be robust enough to work with "cold-start" users who lack extensive social histories.

Takeaway: In the fight against online toxicity, the company you keep speaks as loudly as the words you type.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Neural Networks (GNNs) instead of Node2Vec for social-context-aware hate speech detection in multilingual settings.
  • What are the ethical implications and privacy concerns discussed in literature regarding the use of user follower/retweet graphs for profiling social media users?
  • Find research that applies the concept of Homophily to detect targeted harassment or cyberbullying in low-resource non-English languages.
Contents
Beyond Words: Leveraging Social Milieu for Arabic Hate Speech Detection
1. TL;DR
2. The Problem: The "Hostile Term" Trap
3. Methodology: Fusing Language and Social Graphs
3.1. 1. Dual-Track Word Embeddings
3.2. 2. Social Network Embeddings (The "Milieu")
4. Experimental Results & Insights
4.1. Key Observations:
5. Final Thoughts & Future Directions