NERUD: Capturing the "Language of Doubt" for Robust Rumor Detection
A Neural Rumor Detection Framework by Incorporating Uncertainty Attention on Social Media Texts
NERUD is a neural rumor detection framework that integrates dual attention mechanisms to capture both event semantics and uncertainty expressions. By leveraging a dedicated Uncertainty Identifier alongside a Rumor Detector, it achieves state-of-the-art performance on the Chinese Rumor Corpus (CRC) and demonstrates superior robustness across changing topics.
TL;DR
Rumors spread faster than the truth, often because they exploit public anxiety. While most AI models try to detect rumors by looking at what is being discussed (event semantics), NERUD (Neural Rumor Detection) introduces a shift in perspective. It focuses on how it's being discussed by incorporating Uncertainty Attention. This allows the model to stay effective even when rumors move from earthquake scares to financial panics.
Problem: The Fragility of Topical Semantics
Existing rumor detection systems are often "topic-bound." They learn that certain keywords (like "nuclear radiation" or "stock market crash") are associated with rumors in a specific dataset. However, as topics change, these models lose their edge.
The authors identify a critical missing piece: Uncertainty. According to the classic definition, a rumor is an unverified statement. The language used to convey such statements often includes specific markers of doubt—"witnessed," "probably," "it is said." Unlike specific event words, these uncertainty expressions are consistent across different rumors.
Methodology: The Dual-Attention Framework
NERUD consists of two main modules working in tandem: the Rumor Detector and the Uncertainty Identifier.
1. Hierarchical Post & Event Encoding
The model uses a Bi-GRU (Bidirectional Gated Recurrent Unit) to encode individual posts, followed by another GRU to encode the entire event (the sequence of posts).
2. The Uncertainty Attention Mechanism
This is the "secret sauce." The Uncertainty Identifier is trained to recognize expressions of doubt. The attention weights from this module () are then fused with the standard rumor attention weights ().

As shown in the architecture above, the fusion happens at two levels:
- Post Level: Combining topical importance with uncertainty importance to form a weighted post representation.
- Event Level: Using the uncertainty probability distribution of posts to weight their contribution to the final event representation.
Experiments & Visual Evidence
The authors tested NERUD on Ma's Dataset (Twitter/Sina Weibo) and a new Chinese Rumor Corpus (CRC).
Performance Gains
In the CRC dataset, which tests the model's ability to handle diverse topics, NERUD reached an F1-score of 0.896, beating the previous SOTA model, CSI (0.867), and a version of itself without uncertainty features (NERUD*, 0.858).
| Method | Precision | Recall | F1-Score |
|---|---|---|---|
| ML-GRU | 0.819 | 0.880 | 0.848 |
| CSI | 0.831 | 0.906 | 0.867 |
| NERUD | 0.872 | 0.921 | 0.896 |
The Heatmap: Rumor vs. Uncertainty
The most compelling evidence for "Why it works" is found in the attention visualization.

Fig: Ru Att (Rumor Attention) vs. Un Att (Uncertainty Attention).
In the example "Salt could probably prevent radiation":
- Rumor Attention focuses on "Salt" and "Radiation" (The Event).
- Uncertainty Attention focuses on "Could" and "Probably" (The Doubt).
- Combined Attention provides a holistic representation that marks the post as highly suspicious.
Critical Insight & Conclusion
The success of NERUD demonstrates that rumor detection is not just a classification task based on content, but a stylistic analysis task. By explicitly modeling the context of uncertainty, the framework becomes less susceptible to the "overfitting" of topical trends that plagues traditional NLP models in this domain.
Future Outlook: While NERUD is powerful for text, rumors on platforms today are increasingly multimodal. A logical next step for this research would be applying "Uncertainty Attention" to discrepancies between images and their captions.
