AuthCrowd: Resolving Academic Identity Crisis through Professional Crowdsourcing
AuthCrowd: Author Name Disambiguation and Entity Matching using Crowdsourcing
AuthCrowd is a crowdsourcing system designed for Author Name Disambiguation (AND) and entity matching in bibliometric datasets. It integrates automated data acquisition with a human-in-the-loop framework, achieving over 80% accuracy in complex metadata transcription tasks.
TL;DR
The exponential growth of scientific literature has made Author Name Disambiguation (AND) a nightmare for digital libraries. AuthCrowd introduces a specialized crowdsourcing framework that leverages human intuition to solve the "homonym/synonym" problem in bibliometrics. By decomposing complex matching into digestible micro-tasks, the system achieves over 80% accuracy in metadata validation, proving that the human-in-the-loop approach is vital for high-quality scientometric data.
The "Identity Crisis" in Scientometrics
In the world of big data, "John Smith" is more than a name; it is a point of failure. Modern algorithms struggle when:
- An author changes affiliations (Temporal Ambiguity).
- Multiple authors share near-identical names (Homonymy).
- Names are misspelled or inconsistently initialized (Synonymy).
While Unsupervised Learning (Clustering) and Network Embeddings are the current SOTA, they are brittle when faced with missing data. The authors of AuthCrowd argue that we have hit a plateau with pure ML and must return to Human Intelligence—specifically, a structured, scalable "Crowd" approach—to clean the metadata registry.
Methodology: The AuthCrowd Architecture
The system treats disambiguation not as a single calculation, but as a multi-stage pipeline:
- Task Design: Presenting publications in varied visual formats (e.g., Name and Affiliation in Same Row (NAR) vs. Separated (NAS)) to minimize cognitive load.
- Validation: Using "Ground Truth" questions to filter out low-quality contributors.
- Aggregation: Employing techniques like Majority Voting and HoneyPot to distill a single, high-confidence answer from diverse crowd inputs.
Fig 1: The AuthCrowd interface demonstrating side-by-side article comparison for entity matching.
Experimental Insights: Does the Crowd Deliver?
The study evaluated parameters like Visual Representation and Confidence Scoring.
Key Findings:
- The Difficulty Gradient: Tasks were categorized from A (Hard - different affiliations) to C (Easy). The crowd performed best when the task involved identifying a "Corresponding Author," reaching SOTA-level precision.
- Behavioral Efficiency: Log analysis showed a "learning curve." As users became familiar with the interface, the "Average Number of Clicks" decreased while accuracy remained stable or improved.
- The Power of Confidence: There was a strong correlation between a participant's self-reported "Confidence Level" and the actual "Correctness" of the result. High-confidence positive answers (Scale 4-5) yielded over 75% accuracy.
Table 1: Overview of task complexity vs. accuracy outcomes.
Why This Matters: Moving Beyond the Algorithm
The core contribution of AuthCrowd is the demonstration that crowdsourcing is a "Reasonable Method" for bibliometric IR. Purely algorithmic approaches often treat data points as static; AuthCrowd treats them as part of a narrative that human researchers can navigate more effectively—such as noticing that a change in affiliation is consistent with a typical career progression.
Limitations & Future Work
- Scalability: While 24 participants provided a proof-of-concept, scaling to millions of papers requires a "Hybrid Crowd-Algorithm" strategy where ML handles the 90% easy cases, and AuthCrowd handles the 10% high-entropy edge cases.
- Interface Overlap: User feedback suggested that side-by-side zoom features occasionally caused text overlap, suggesting a need for more responsive UI design.
Conclusion
AuthCrowd isn't just a tool; it's a paradigm shift. It suggests that the future of authoritative scientific databases depends on a Collaborative Intelligence model, where the expert crowd acts as the ultimate supervisor for the AI.
