Expertise-Aware Taxonomy Enrichment: Scaling Human Intelligence via Graph Gaussian Processes

Expertise-Aware Crowdsourcing Taxonomy Enrichment

2021-01-01
Yuquan Wang, Yanpeng Wang, Yiming Mao, Jifan Yu, Kaisheng Zeng, Lei Hou, Juanzi Li, Jie Tang
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces an expertise-aware crowdsourcing framework for Taxonomy Enrichment, designed to classify new instances into thousands of concepts. It leverages a Graph Gaussian Process model to estimate worker skills and includes a subtree recommendation mechanism, achieving a 6.6% accuracy improvement over the BTSK baseline on a 632-node taxonomy.

TL;DR

Taxonomy Enrichment—the process of slotting new instances into a massive hierarchy of concepts—is notoriously difficult for crowdsourced workers. This paper introduces an expertise-aware framework that uses Graph Gaussian Processes to predict worker skills across thousands of nodes based on just a few labels. By matching workers' specific "knowledge pockets" to tasks and providing Subtree Recommendations, the system improves accuracy by 6.6% and reduces annotation time by 7.1%.

The Problem: The "Curse of Knowledge" in Crowdsourcing

Taxonomies like YAGO or MOOCCube contain hundreds of thousands of concepts. For a human worker, finding the exact leaf node for a term like "alligator" (Is it a Reptile? An Amphibian? Which specific sub-branch?) is a daunting task.

Traditional crowdsourcing falls short in two ways:

  1. The Quiz Bottleneck: You can't give a worker a test on 10,000 concepts to see if they are an expert.
  2. The Search Burden: Forcing workers to navigate the entire tree or answering endless binary "Yes/No" questions is exhausting and kills motivation.

The authors' core insight is Skill Locality: if a worker knows about "Amphibians," they likely know about "Reptiles," but might know nothing about "Quantum Physics." Expertise is not random; it is clustered on the taxonomy graph.

Methodology: Skill Modeling and Recommendations

The framework operates in a continuous loop of Task Assignment and Answer Aggregation.

1. Skill Estimation via Graph Gaussian Process (GGP)

The system models worker expertise as a latent variable over the taxonomy graph. Using a Radial Basis Function (RBF) kernel based on the graph distance (number of edges), the model propagates known performance data to neighboring nodes.

  • Accuracy Locality: Workers have similar performance on nearby concepts.
  • Bias Locality: If a worker makes a mistake, they likely choose a "neighboring" concept rather than a completely unrelated one.

Overall Framework Architecture

2. Subtree Recommendation

To solve the search burden, the system calculates the probability that an instance belongs to a specific branch. If the probability exceeds a threshold , it presents the worker with that specific subtree rather than the root. This reduces "UI friction" while allowing the worker the autonomy to make final, fine-grained decisions.

Experiments: Superior Quality and Efficiency

The authors tested their approach against BTSK (Budgeted Task Scheduling for Knowledge Acquisition) using both synthetic data and a real-world taxonomy of Science concepts.

Key Findings:

  • Accuracy Boost: The method achieves a significant edge in accuracy (6.6% better than BTSK) by effectively aggregating answers while accounting for worker-specific confusion matrices.
  • Workload Reduction: As shown in the performance curves, the proposed task assignment selects the "most uncertain" instances and matches them to "most expert" workers, reaching high accuracy with far fewer total labels.

Experimental Results Comparison Figure: The proposed method (OTA/OTI) shows a steeper learning curve, reaching higher accuracy with lower label redundancy compared to baselines.

Annotation Time:

By implementing Subtree Recommendations, the average time per label dropped from ~43 seconds to ~40 seconds. While 3 seconds sounds small, across 510,000 instances in the MOOCCube dataset, this represents thousands of hours saved.

Critical Insight & Conclusion

The brilliance of this paper lies in treating Human Expertise as a Manifold. Instead of viewing a worker as a "black box" with a single accuracy score, the authors recognize that human knowledge has a topology.

Limitations: The model relies on a fixed taxonomy structure. If the taxonomy itself is poorly designed or has logical "leaps" (nodes that are semantically close but far in the graph), the Gaussian Process kernels may fail to capture the true skill locality.

In the era of LLMs, this framework remains highly relevant. It provides a blueprint for how we might "curate" LLM outputs using human-in-the-loop systems that don't overwhelm the human, but rather treat them as a surgical expert in a specific domain.

Find Similar Papers

Try Our Examples

  • Find recent papers on Taxonomy Enrichment that utilize Large Language Models (LLMs) as either automated annotators or as a prior for human crowdsourcing.
  • Which original research established the Graph Gaussian Process, and how does this paper adapt that theory to model worker "confusion matrices" in hierarchical labels?
  • Explore how the "subtree recommendation" and expertise-matching logic could be applied to hierarchical classification tasks in medical imaging or legal document labeling.
Contents
Expertise-Aware Taxonomy Enrichment: Scaling Human Intelligence via Graph Gaussian Processes
1. TL;DR
2. The Problem: The "Curse of Knowledge" in Crowdsourcing
3. Methodology: Skill Modeling and Recommendations
3.1. 1. Skill Estimation via Graph Gaussian Process (GGP)
3.2. 2. Subtree Recommendation
4. Experiments: Superior Quality and Efficiency
4.1. Key Findings:
4.2. Annotation Time:
5. Critical Insight & Conclusion