Scaling the Crowd: The Science of Quality Control in FamilySearch Indexing
Quality control mechanisms for crowdsourcing: peer review, arbitration, & expertise at familysearch indexing
This paper evaluates quality control mechanisms for large-scale historical document transcription within FamilySearch Indexing (FSI). It proposes and tests a "Peer Review" (A-R) model as a more efficient alternative to the traditional "Arbitration" (A-B-ARB) method, achieving comparable accuracy while significantly reducing volunteer time.
TL;DR
How do you transcribe billions of handwritten historical records without losing accuracy? FamilySearch Indexing (FSI) proves that while "Double-Entry + Arbitration" is the gold standard for quality, a specialized Peer Review mechanism can deliver 90%+ efficacy while being 30% faster. This study provides a rare, large-scale look at how volunteer expertise and process design dictate the success of global crowdsourcing.
The Bottleneck of Perfection
In the world of genealogy, a single misspelled "Surname" can make a record unsearchable forever. Historically, FSI used the A-B-ARB model:
- Volunteer A transcribes the image.
- Volunteer B independently transcribes the same image.
- An Arbitrator settles the differences.
This ensures high quality but consumes massive amounts of human labor. As the backlog of digitized images grows into the billions, this redundancy becomes a liability. The authors ask: Can we move to a Peer Review (A-R) model (where one person transcribes and another simply corrects) without the quality falling off a cliff?
Methodology: Experience & Flow
The researchers utilized two datasets:
- Historical Data: Over 400,000 volunteers categorized into 9 "Experience Levels."
- Field Experiment: A random sample of 2,000 pages from the 1930 US Census, benchmarked against a 99.75% accurate "truth set."
Logic of the Heatmap
The study analyzed agreement between volunteers. Below is the mental model: if two novices (Level 0) work together, they are far more likely to agree on a wrong answer or disagree entirely.
Figure: The heatmap reveals that agreement (and thus accuracy) scales significantly as both transcribers gain experience (Level 8).
Key Insights: Why Peer Review Wins (Usually)
1. The Efficiency Gap
Peer review is fundamentally faster because identifying an error is cognitively easier than generating a transcription from scratch. The study found the A-R process takes roughly 70% of the time of the traditional A-B-ARB model.
2. The Experience Multiplier
The data shows that expertise isn't just about accuracy; it's about speed. A Level 8 volunteer is 4 times faster than a novice.
- Novice: ~65 seconds per line.
- Expert: ~15 seconds per line.
3. The Quality Trade-off
For "closed" fields (Gender, Race, Age), Peer Review is virtually indistinguishable from Arbitration. However, for "open" fields (Surnames, Given Names), Arbitration is still superior because it prevents the reviewer from "lazy agreement" with the first transcriber.
Table: Comparison of accuracy across different models. Note that A-B-ARB leads in Surnames, but A-R is nearly identical in structured fields.
Designing the Future of Human Computation
The authors suggest several "Heavyweight Peer Production" strategies:
- Intelligent Routing: Route difficult fields (like cursive surnames) to experts, while novices handle "easy" structured fields (like Age).
- Contextual Learning: Volunteers perform better when they have domain knowledge. Someone transcribing Canadian records should be served Canadian images consistently to build a "mental lexicon" of local placenames.
- Removing Redundant Verification: Surprisingly, adding a third step to Peer Review (A-R-RARB) actually lowered accuracy. Reviewers were already conservative and only changed things when they were sure; adding another layer just introduced more "noise."
Conclusion
The study concludes that the "Golden Age" of redundant crowdsourcing is evolving. By moving toward a Peer Review model and leveraging the disproportionate impact of Expert Volunteers, platforms like FamilySearch can process history 30% faster without sacrificing the integrity of the data. For researchers, the takeaway is clear: optimize for expertise and flow, not just redundant "eyeballs."
