When Humans Forget: Designing Bias-Aware Systems for Social Stream Processing

Modeling human annotation errors to design bias-aware systems for social stream processing

2019-08-27
Rahul Pandey, Carlos Castillo, Hemant Purohit
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a bias-aware active learning framework designed for social media stream processing. It proposes a novel "Error-mitigating Sampling" method that models human forgetting behavior using the Ebbinghaus Curve to optimize the annotation schedule, effectively improving classification AUC in environments prone to human errors.

TL;DR

High-quality machine learning requires high-quality human data. But humans aren't perfect—we forget, get tired, and make mistakes. This paper introduces a Bias-Aware Hybrid System that models human forgetting behavior (based on the Ebbinghaus Curve) to decide which data points to show annotators. By optimizing the annotation schedule, the system boosts classification accuracy in social media crisis monitoring by mitigating human error before it happens.

Problem & Motivation: The "Burnout" of the Human Oracle

In the middle of a disaster like Hurricane Harvey, social media streams are a firehose of information. Hybrid Stream Processing Systems (HSPS) rely on humans to label "infrastructure damage" or "rescue efforts" to train AI models in real-time.

The industry's blind spot has long been the assumption that human labels are "Ground Truth." In reality:

  • Burnout: High cognitive load leads to deteriorating quality.
  • Slips & Mistakes: Forgetting a category exists simply because you haven't seen it in the last 100 tweets.
  • The Schedule Gap: Conventional active learning picks the most "uncertain" data for humans but ignores whether the human is in a fit state (cognitively) to label that specific instance correctly.

Methodology: Modeling the Human Mind

The researchers' core insight is that human error isn't random; it's predictable.

1. The Forgetting Curve

They utilize the Ebbinghaus Curve, modeled as a sigmoid function, to calculate a forgetting_score. If a class (e.g., "Utility Damage") hasn't appeared in the stream for a while, the probability of a human "slipping" increases.

Ebbinghaus Forgetting Curve

2. Error-Mitigating Sampling

Standard Active Learning (Uncertainty Sampling) focuses on the model's confusion. The proposed Algorithm 3 (Error-mitigating Sampling) adds a second filter:

  1. Predictive Uncertainty: Identify samples where the model is 30-70% confident.
  2. Bias & Forget Score: Calculate if showing this specific class right now will induce a performance error or if the human has likely forgotten the class representation.
  3. Discard & Update: Discard instances that are likely to result in erroneous human labels, thereby protecting the model from learning from "poisoned" data.

Experiments & Results

The team tested their approach against two hurricanes (Harvey and Irma) using Twitter data. They simulated three types of oracles: "Slow Forgetting," "Fast Forgetting," and "No Forgetting."

Key Findings:

  • Superior Robustness: Even with "Fast Forgetting" (highly error-prone humans), the error-mitigating algorithm maintained a stable and higher AUC compared to random sampling.
  • Learning Curve: While simple uncertainty sampling works initially, its performance fluctuates wildly as human errors accumulate. The proposed method acts as a stabilizer.

Performance Comparison Fig: AUC scores across datasets. The error-mitigating sampling (Proposed) shows more consistent growth and higher peaks than standard baselines.

Critical Analysis & Conclusion

Takeaway

This paper shifts the focus of Active Learning from "What does the model need to know?" to "What is the human capable of teaching right now?" By treating the human annotator as a dynamic, error-prone component of the system rather than a static oracle, we can design much more resilient AI.

Limitations

  • Text-Only: The current study focuses on Twitter text; cognitive load might behave differently for image or multi-modal tasks.
  • Parameter Sensitivity: The sigmoid function requires specific parameters () which might vary significantly between different crowdsourcing pools.

Future Work

The next frontier is extending this to Human-AI Collaboration where the system provides "hints" (e.g., showing a reference image of the class) to "refresh" the human's memory when the forget_score gets too high.

Find Similar Papers

Try Our Examples

  • Identify recent papers that integrate cognitive psychology models, such as the Ebbinghaus forgetting curve, into human-in-the-loop machine learning workflows.
  • Who first defined the taxonomy of "slips" and "mistakes" in human error research, and how has this been adapted for modern crowdsourcing tasks?
  • Examine how current SOTA online active learning methods for concept drift mitigation handle noisy or sub-optimal oracle labels in real-time environments.
Contents
When Humans Forget: Designing Bias-Aware Systems for Social Stream Processing
1. TL;DR
2. Problem & Motivation: The "Burnout" of the Human Oracle
3. Methodology: Modeling the Human Mind
3.1. 1. The Forgetting Curve
3.2. 2. Error-Mitigating Sampling
4. Experiments & Results
4.1. Key Findings:
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Work