Scaling the Library: How Crowdsourcing Reinvigorated the INEX 2010 Book Track

Overview of the INEX 2010 Book Track: Scaling Up the Evaluation Using Crowdsourcing

2011-01-01
Gabriella Kazai, Marijn Koolen, Jaap Kamps, Antoine Doucet, Monica Landoni
Summary
Problem
Method
Results
Takeaways
Abstract

This paper provides an overview of the INEX 2010 Book Track, which evaluates information retrieval (IR) and structure extraction methods for a corpus of 50,000 digitized books. Its primary contribution is the successful integration of Amazon Mechanical Turk (AMT) crowdsourcing to scale up the creation of test topics and relevance judgments, achieving high alignment with professional editorial standards.

TL;DR

Evaluation at scale has long been the "Achilles' heel" of Information Retrieval (IR) research. The INEX 2010 Book Track addressed this by pivoting from purely expert-led evaluations to a crowdsourced-hybrid model. By leveraging Amazon Mechanical Turk (AMT), the organizers scaled a corpus of 50k books, achieving a 78% agreement rate with experts and validating that "the crowd" can indeed handle complex academic retrieval tasks.

The Bottleneck: Why Expert Evaluation Failed to Scale

In the quest to improve how we search and navigate digitized libraries (like Google Books or the Million Book Project), researchers hit a wall. Using the traditional Cranfield method—where experts manually judge every result for relevance—was no longer sustainable.

In 2008, it was estimated that a single participant would need to spend over an hour a day for a full month just to judge one topic. This "daunting" workload led to "participation churn": many labs registered, but few actually submitted results. To save the track, the organizers had to find a way to "Scale Up" without sacrificing the academic rigor of the test collection.

Methodology: Designing for the Crowd

The 2010 track focused on four tasks, with Best Books to Reference (BB) and Prove It (PI) being the primary search challenges. The innovation lay in how these tasks were benchmarked:

  1. Topic Creation: Workers were paid to find "facts" that appeared both in a book and on Wikipedia, ensuring topics were grounded in real-world information needs.
  2. Quality Control: The organizers didn't just trust the crowd blindly. They used a "Trusted vs. Crowd" hybrid approach:
    • Worker Pre-selection: Only workers with a >95% approval rate were allowed.
    • Gold Set Injection: Each task (HIT) included at least one page already labeled by a "trusted" INEX expert to check for worker accuracy.
    • Majority Voting: Three workers judged every page to reach a consensus.

Relevance Assessment Module Figure 1: The interface used by both experts and workers to judge book relevance.

Experiments & SOTA Results

The track evaluated several sophisticated approaches, ranging from language modeling to page-level logistic regression.

  • Best Books (BB) Results: The University of California, Berkeley, took the lead (NDCG@10 = 0.6579) by summing individual page scores to derive a total book relevance score.
  • Prove It (PI) Results: The University of Amsterdam excelled (NDCG@10 = 0.2946) by using individual page-level indexing combined with Pseudo-Relevance Feedback (PRF).

The "Prove It" task is particularly difficult—it is a "needle-in-a-haystack" problem where systems must find specific pages that confirm or refute a factual claim. To account for this, the organizers measured "Near-Misses," showing that many systems found relevant information in the vicinity of the target page even if they missed the exact XPath.

Prove It Task Results Figure 2: Evaluation results for the Prove It task, showcasing performance when accounting for proximity-based 'near-misses'.

Deep Insight: Is the Crowd Reliable?

The most critical finding was the 78% agreement rate between AMT workers and INEX experts for binary relevance. While workers were slightly more likely to label a page as "confirm/refute" than experts, the consensus (majority vote) was remarkably stable.

This suggests that for most IR tasks, the "Editorial Bottleneck" is a choice, not a necessity. By decentralizing the labor, the INEX 2010 track successfully built a larger, more diverse test collection than was ever possible with academics alone.

Conclusion and Future Outlook

The 2010 track didn't just find the "Best Books"; it found a Best Practice for the future of IR evaluation. By moving toward crowdsourcing, the track transitioned from a struggling niche project to a scalable framework.

Future directions identified in the paper include moving toward "Social Search"—incorporating user-generated metadata like reviews and tags (from Amazon and LibraryThing)—further blurring the line between professional library science and the collective intelligence of the web.


Key Takeaway: Crowdsourcing, when paired with defensive task design and "Gold Set" validation, is an academically viable tool for scaling IR evaluations to the massive datasets of the digital age.

Find Similar Papers

Try Our Examples

  • Find recent studies that compare the cost-effectiveness and accuracy of crowdsourced relevance judgments versus expert editorial assessments in large-scale IR benchmarks.
  • Which paper first established the 'Gold Set' methodology for quality control in human computation, and how has it evolved for complex semantic tasks?
  • Explore how crowdsourcing and active learning are currently being combined to build ground-truth datasets for OCR structure extraction and document layout analysis.
Contents
Scaling the Library: How Crowdsourcing Reinvigorated the INEX 2010 Book Track
1. TL;DR
2. The Bottleneck: Why Expert Evaluation Failed to Scale
3. Methodology: Designing for the Crowd
4. Experiments & SOTA Results
5. Deep Insight: Is the Crowd Reliable?
6. Conclusion and Future Outlook