Active Learning with Confidence: Beyond Binary Labels in Crowdsourcing
KNOWLEDGE‐BASED SYSTEMS
The paper introduces a novel active learning framework for crowdsourcing that utilizes confidence-based (continuous) answers rather than traditional discrete labels. It proposes a Maximum Likelihood Estimation (MLE) aggregation method based on the Beta distribution and a new instance selection score that combines model uncertainty with answer set uncertainty.
TL;DR
In crowdsourcing, a "Yes" isn't always a "Yes"—sometimes it's a "Maybe." This paper proposes a paradigm shift from discrete labeling to confidence-based continuous responses. By modeling worker uncertainty through a Beta distribution and integrating this into an Active Learning (AL) framework, the authors achieve SOTA accuracy on difficult datasets while significantly reducing labeling costs.
Problem: The "Illusion of Certainty" in Crowdsourcing
Most crowdsourcing systems treat workers as noisy oracles providing discrete labels (e.g., Cat vs. Dog). However, for borderline cases—like a person who looks exactly 20 years old in an "Under 20?" task—forcing a binary choice loses vital information.
Previous SOTA methods attempted to solve this by:
- Modeling Worker Reliability: Tracking history to weight experts (fails for anonymous, transient workers).
- Repeated Labeling: Hiring more workers until a consensus is reached (wasteful and expensive).
The authors argue that the problem isn't just who is labeling, but what they are providing.
Methodology: Probability as a First-Class Citizen
The core innovation lies in two areas: how labels are collected and how they are selected.
1. The Beta Distribution Aggregator
Workers use a slider to estimate . The researchers assume these probabilities follow a Beta distribution, . Using a reparameterization involving the mean and a precision parameter , they apply Maximum Likelihood Estimation (MLE) via the BFGS algorithm to find the most likely "true" probability.
Fig 1. The iterative Active Learning framework leveraging continuous worker feedback.
2. Doubly-Uncertain Score Function
When selecting the next instance to label, the system doesn't just look at what the model is "confused" about (Model Uncertainty). It also looks at how "confused" the current set of human labels is (Answer Uncertainty).
The Answer Uncertainty () is defined by the Confidence Interval (CI) of the mean. If the CI contains 0.5 (neutrality), the instance is deemed highly uncertain.
Experimental Proof: Winning at Difficult Tasks
The paper rigorously tests the framework on "difficult" versions of standard datasets (Mushroom, Bank, Tic-tac-toe) where instance labels are naturally ambiguous.
Key Findings:
- Higher Accuracy: On the FG-NET face dataset, the proposed HIS5 method outperformed baseline active learners by 10% in challenging age-identification tasks.
- Adaptive Budgeting: The model intelligently allocated more budget (more workers per image) to "20-ish" faces—the ones humans find hardest to judge—while spending less on obviously young or old faces.
Fig 2. Accuracy comparison: The MLE aggregation method (red line) consistently outperforms traditional Majority Voting.
Efficiency at Scale: Batch Selection
To make this practical for real-world systems where retraining a model for every single label is too slow, the authors introduced Top-k and Clustering-based batch selection. Clustering ensures the selected batch is diverse, preventing the model from wasting the budget on 50 identical-looking images. This saved over 95% of computational time with negligible loss in accuracy.
Critical Insight & Conclusion
This work demonstrates that soft-labeling isn't just a UI gimmick; it's a mathematically superior way to handle aleatoric uncertainty in data. By allowing workers to "express" their doubt, the system converts that doubt into a signal for the Active Learner to either hire more workers or move on to more informative samples.
Future Outlook: While highly effective for binary tasks, moving to multi-class scenarios is the next frontier. How do we collect confidence across 1,000 ImageNet classes without overwhelming the user? The authors suggest using sparse distributions as a potential solution.
