Breaking the NIST Bottleneck: Can the Crowd Replace Expert IR Assessors?
Identifying top news using crowdsourcing
This paper investigates the feasibility of using crowdsourcing via Amazon Mechanical Turk to generate relevance assessments for the TREC 2010 Blog track. It introduces a large-scale evaluation framework for "News Story Ranking" and "Blog Post Ranking" tasks, demonstrating that crowdsourcing can serve as a viable, cost-effective substitute for traditional expert assessors.
TL;DR
This seminal study from the University of Glasgow proves that crowdsourcing is not just for simple image tagging—it is a robust, "industrial-strength" solution for complex Information Retrieval evaluation. By generating over 30,000 judgements for the TREC 2010 Blog track, the authors demonstrated that non-experts, managed with the right redundancy and validation, can produce SOTA-level ground truth data at 1/3 the cost of traditional experts.
Background: The Scalability Crisis in IR
For decades, the Text REtrieval Conference (TREC) has been the gold standard for evaluating search engines. However, its reliance on specialized assessors creates a bottleneck: it is expensive and slow. As datasets ballooned into the millions of documents, the "pooling" method became increasingly incomplete. The core question for McCreadie et al. was: Can we trust anonymous workers on Mechanical Turk to judge the "newsworthiness" of a blog post or the "importance" of a news story?
Methodology: Designing for the Crowd
The authors tackled two distinct sub-tasks with different technical "inductive biases" in their HIT (Human Intelligence Task) design:
1. News Story Ranking (Relative Context)
Judging if a news story is "top tier" requires knowing other stories from the same day.
- The Solution: Large HITs spanning 32 judgements. This forced workers to maintain context rather than judging in a vacuum.
- Visual Validation: Instead of just math-based gold standards, they used color-coded summary interfaces to detect "bot-like" behavior at a glance.

2. Blog Post Ranking (Cleaned Content)
Blog posts are messy, filled with dead CSS and broken ads.
- The Solution: A "Cleaned Document" view. By stripping HTML to basic text, they reduced loading times and worker cognitive load, resulting in a 10x speedup in completion rates.
Key Insights: The Efficiency of Redundancy
One of the paper’s most vital contributions is the evaluation of the "2+1" Strategy.
- In a standard 3-worker redundancy, you pay for 3 judgements every time.
- In the "2+1" approach, you only hire a third worker if the first two disagree.
- The Result: 26.7% cost savings with zero change in the final aggregated ground truth. This is a critical finding for researchers with limited budgets.
Experimental Results: Accuracy and Agreement
The "crowd" proved surprisingly reliable. Inter-worker agreement (Fleiss Kappa) reached 83% for objective categories like "Science/Technology" and "Sports," though it dipped for the more subjective "World News" category where regional bias might occur.

Does it change the leaderboard? The authors performed a "ranking stability" test. They found that while single-assessor data changed the leaderboard of TREC systems, the majority-vote (redundant) crowd data effectively matched the discrimination power of expert-led evaluations.
Critical Analysis & Best Practices
The paper concludes with a "Manual of Best Practices" that remains relevant:
- Don’t Fear Large HITs: Consolidating 20+ judgements per HIT reduces context-switching and improves worker retention.
- External Integration: Use your own judging software via IFrames (ExternalQuestion) to log fine-grained worker behavior.
- Active Re-costing: As workers get faster (the "learning curve"), adjust payments to ensure fair but efficient allocation of funds.
Limitations
While the study proved feasibility, it also noted that crowdsourcing isn't always "instant." The blog post task took two weeks to complete, highlighting that task complexity and worker availability (e.g., timezone differences) still play a major role in the evaluation pipeline.
Conclusion
This work was the first to successfully "crowdsource TREC" at scale. It shifted the perspective from viewing Mechanical Turk as a "noisy source" to seeing it as a configurable, high-throughput engine for scientific evaluation. Today, as we transition to LLM-based evaluation, the principles of redundancy and systematic validation laid out in this paper remain the fundamental bedrock of IR experimentation.
