Beyond Simple Voting: Aggregating Crowdsourced Labels via Worker History
Aggregation of Crowdsourced Labels Based on Worker History
The paper introduces a novel EM-based framework for aggregating crowdsourced binary labels by iteratively estimating worker expertise (confidence) and the true item labels. The method achieves significant performance gains over Majority Voting, notably reaching approximately 93% F1 score on the RTE-RTE dataset.
TL;DR
Crowdsourcing is cost-effective but notoriously noisy. This paper presents an Expectation-Maximization (EM) framework that moves beyond Majority Voting by modeling worker expertise as a dynamic "confidence" score. By looking at a worker's history and weighing their votes based on their reliability—and optionally their self-reported familiarity—this method significantly cleans up noisy datasets, outperforming established baselines like ZenCrowd and Dawid-Skene on diverse benchmarks.
Background: The Limits of Democracy in Crowdsourcing
In supervised machine learning, the quality of training data is the "glass ceiling" of model performance. While platforms like Amazon Mechanical Turk offer scale, they also introduce "spammers" or non-experts. The traditional solution, Majority Voting (MV), treats every vote as equal. However, MV fails when:
- Workers have varying skill levels: A specialist's vote should count more than a novice's.
- Tasks vary in difficulty: Some items are inherently ambiguous, leading to ties that MV cannot resolve intelligently.
Methodology: The Mutual Reinforcement Loop
The core insight of the authors is that item labels and worker confidence are two sides of the same coin. If we knew the true labels, we could identify the best workers; if we knew the best workers, we could find the true labels.
The authors solve this "chicken and egg" problem using an EM Algorithm:
- E-Step (Aggregation): Compute "soft" labels () for each item. This isn't just a count; it's a weighted sum where each worker's vote is multiplied by their confidence score.
- M-Step (Update): Update the worker's confidence () based on how well their historical votes align with the aggregated labels produced in the E-step.
Key Framework Enhancements
- Soft Evaluation: Unlike hard nominal updates, soft evaluation uses the "strength" of the crowd's agreement. If a worker agrees with a highly certain crowd, their confidence increases more.
- PN-Discrimination: The model tests whether workers are better at identifying "positives" versus "negatives"—a common phenomenon in binary classification tasks.
- Confidence Boosting: Applying non-linear functions (like or ) to the confidence scores to further separate experts from average workers.
(Note: The logic above reflects the asymmetric confidence update for Positive/Negative discrimination)
The "Familiarity" Factor
A unique contribution of this research is the inclusion of Self-Reported Familiarity. The authors found that when workers claim low familiarity, they are statistically more accurate at giving negative answers than positive ones. By adjusting worker weights based on these subjective self-assessments, the model gains an extra layer of inductive bias that purely statistical models lack.
Experimental Performance
The method was tested against the SQUARE Benchmark, including diverse datasets like RTE (Textual Entailment) and WVSCM (Smile recognition).

- RTE-RTE: Achieved a massive boost from 0.89 (MV) to 0.93 (Proposed).
- WVSCM: While traditional methods like Dawid-Skene and Raykar often struggle with sparse labels, this method showed the most consistent accuracy improvements.
- Fashion Datasets: In tasks requiring domain expertise (identifying clothing styles), incorporating familiarity coupled with soft evaluation proved most effective.
Critical Insight: Why Does It Work?
By allowing Soft Labels, the algorithm captures the uncertainty of a task. In a tie-break situation where three workers say "Yes" and three say "No," Majority Voting flips a coin. This EM approach looks at the workers' histories: if the "Yes" voters have historically been more accurate on similar tasks, the tie is broken with high-confidence logic rather than randomness.
Conclusion & Future Outlook
This work demonstrates that worker history is a goldmine for quality control. While no single configuration (Hard vs. Soft, PN vs. Non-PN) wins across every dataset, the primary framework is robust enough to provide a superior "Ground Truth" for training downstream ML models.
Future research could investigate how to apply these confidence-weighting strategies to Generative AI evaluations, where "labels" are complex textual responses rather than simple binary choices.
