Is the Crowd an Assistant or a Replacement? Navigating Expert-Level Ontology Engineering
Is the crowd better as an assistant or a replacement in ontology engineering? An exploration through the lens of the Gene Ontology
This study evaluates a crowdsourcing-based methodology for verifying the Gene Ontology (GO), following success in clinical ontologies like SNOMED CT. Using platforms like Mechanical Turk and CrowdFlower, the research explores whether the crowd can act as a replacement for or an assistant to domain experts in identifying ontological errors.
TL;DR
Can a group of non-experts on the internet find errors in the highly specialized Gene Ontology (GO)? While previous work suggested the "wisdom of the crowd" could match medical experts in clinical settings (SNOMED CT), this study finds that when the knowledge becomes too esoteric, the crowd falters. However, by using "Google-ability" as a metric, we can identify "easy" tasks the crowd can handle, transforming them from expert replacements into high-efficiency assistants in a "group-sourcing" workflow.
Background: The Scalability Crisis in Bio-Ontologies
Biomedical ontologies like the Gene Ontology are essential for data integration and enrichment analysis. However, as they grow to include tens of thousands of concepts, maintaining their accuracy becomes a bottleneck. Automated tools often catch structural flaws but miss nuanced semantic errors (e.g., mislabeling a transition metal response). Currently, the "gold standard" is manual expert review—accurate but prohibitively expensive and slow.
The "Why": From Clinical Practice to Molecular Biology
The authors previously demonstrated that crowd workers could verify SNOMED CT (clinical terms) with expert-level accuracy. The core motivation here was to see if this success scaled to the Gene Ontology. GO is more "esoteric"—while a layperson might have an intuition about "Brain Disorders," they likely know nothing about "Extrinsic components of the stromal side of the plastid inner membrane."
Methodology: The Verification Pipeline
The researchers extracted 200 complex, logically entailed relationships from GO. They established a consensus baseline using five experts through a Delphi Method (a systematic, interactive forecasting method relying on a panel of experts).
The crowd was then presented with natural language sentences representing these relationships (e.g., "A is a kind of B") and asked to verify them as True or False based on provided definitions.
Figure 1: Comparison of the task interface presented to experts vs. the crowd.
Why the Crowd Struggled: The "Google-ability" Factor
The results were humbling. Unlike the clinical success in SNOMED CT, the crowd’s performance on GO was highly variable and often poor (AUC as low as 0.44). The study identified three critical factors for success:
- Task Difficulty: Relationships where experts disagreed were also where the crowd failed most.
- Definition Utility: When the provided text definitions were clear and useful, crowd accuracy spiked.
- Search Volume: GO terms are significantly less common on the internet than SNOMED terms.
Figure 2: Empirical cumulative distribution of search results. Note the stark difference between GO and SNOMED CT visibility.
The researchers found that the number of Google search results for a concept is a strong predictor of whether the crowd will succeed. If a term is "Google-able" (like "Response to acid"), the crowd wins. If it is esoteric, they fail.
Conclusion: Toward "Group-Sourcing"
The paper concludes that the crowd is not an expert replacement but a powerful Expert Assistant.
The Takeaway for Developers: Don't send every task to the crowd. Build a "Group-Sourcing" agent that:
- Filters: Uses computational metrics to find potential errors.
- Routes: Checks search result volume. High volume? Send to the crowd. Low volume? Route to an expert.
- Learns: Continually refines which tasks go to whom.
By offloading the "easy" 50% of ontology verification to the crowd, experts can focus their expensive time on the complex biological nuances where they are truly needed.
Figure 3: Performance varies strikingly across different strata of definition utility and expert agreement.
