Measuring Arabic Tweet Credibility: Bridging the Language Gap in Information Verification
Measuring the credibility of Arabic text content in Twitter
This paper introduces an automated tool to evaluate the credibility of Arabic news content on Twitter. It employs two primary methodologies: a similarity-based approach comparing tweets to authoritative news sources (like SPA and Aljazeera) and a multi-feature scoring model incorporating user metadata and external rankings.
TL;DR
In an era of rapid information cycles, this research presents the first automated framework specifically designed to measure the credibility of Arabic news on Twitter. By comparing tweets against verified news agencies and analyzing metadata, the authors developed a system that achieves a 0.52 precision rate, proving that linguistic-aware similarity checks often outperform complex social metrics.
Background & Motivation
Twitter has become a primary source for real-time news, yet it is a double-edged sword prone to gossip and misinformation. While automated credibility tools exist for major languages, the Arabic web has been largely ignored. Arabic presents unique Natural Language Processing (NLP) hurdles:
- Morphological Complexity: Highly inflectional and derivational structures make stemming difficult.
- Lack of Diacritics: Identical-looking words can have entirely different meanings.
- No Capitalization: Identifying proper nouns (like "Saudi Press Agency") requires sophisticated context rather than simple casing rules.
Methodology: Two Paths to Truth
The researchers proposed a system architecture consisting of five phases: Data Input, Preprocessing, Feature Computation, Calculation, and Ranking.
1. The Similarity Approach (Semantic Mapping)
This method creates a "Centroid Vector" from 179 authoritative news articles. Every tweet is converted into a Bag Of Words (BOW) and compared against this centroid using Cosine Similarity.
- Insight: To handle the length disparity (short tweets vs. long articles), they modified the IDF calculation to prevent zero-division when tweet terms were missing from the news corpus.
2. The Multi-Feature Formula
The system also calculates a weighted score based on:
- Content (60% Similarity + 20% Cleanliness + 10% Linking): Checks for inappropriate words and verified external URLs.
- Author (5% Verification + 5% Grader Score): Uses external APIs like TwitterGrader to assess user influence.

Experimental Analysis
The authors collected a dataset of 600 tweets and 179 news articles focusing on hot political topics.
Key Findings:
- The "Nouns" Advantage: Using similarity based only on nouns with light stemming (lemmatization) provided the most consistent results.
- Small Similarity Scores: Average similarity values were below 0.5. This is attributed to the brevity of tweets (~10 words) compared to full-length articles.
- Feature Weighting: Surprisingly, the complex scoring model (Approach 2) often performed worse than simple similarity. The "Linking" feature was the only secondary feature that showed significant impact.

Above: The distribution of credibility levels. Note that similarity alone (Approach 1) struggled to identify "Average" credibility, tending to polarize tweets into High or Low.
Evaluation by Experts
Human experts in political science were used to benchmark the tool.
- Precision: Averaged between 0.45 and 0.55.
- Recall: Averaged between 0.35 and 0.57.
The results indicate that while the tool is a strong starting point, it captures "Low Credibility" (obvious misinformation) much more accurately than "High Credibility."
Critical Insight & Conclusion
The study concludes that for the Arabic language, linguistic preprocessing is the bottleneck. Simple light stemming—focusing specifically on nouns—drastically improves the system's ability to match tweets with authoritative sources.
Future Directions: The authors suggest that future iterations should incorporate Twitter-specific markers like hashtags (#), retweets, and emoticons to refine the "Questionable" category. As social media continues to dominate the Middle Eastern news landscape, such localized credibility tools are no longer optional—they are essential infrastructure for digital literacy.
