Beyond Binary: Managing Crowdsourcing Ambiguity with Interval-Valued Labels

Managing Uncertainty in Crowdsourcing with Interval-Valued Labels

2021-07-28
Chenyi Hu, Victor S. Sheng, Ningning Wu, Xintao Wu
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Interval-Valued Labels (IVLs) for crowdsourcing to manage human-induced ambiguity. By allowing workers to provide range-based responses (e.g., [0.7, 0.8]) instead of binary 0/1 labels, the authors propose a new inference algorithm that leverages the statistical properties of intervals to achieve a matching probability exceeding 50%.

TL;DR

In the world of crowdsourcing, we usually ask workers for a "Yes" or a "No." But human intuition isn't binary—it's fuzzy. This paper proposes Interval-Valued Labels (IVLs), allowing workers to express uncertainty as a range (e.g., "I'm 70% to 80% sure this is a 'Yes'"). This shift from points to intervals prevents information loss and enables a sophisticated probabilistic inference that outperforms traditional Majority Voting.

The Problem: The "Forced Choice" Trap

Crowdsourced data is the lifeblood of modern AI, yet it is notoriously noisy. Current systems typically compress human judgment into binary logic:

  • Physical constraints: Is this a cat? (1 or 0).
  • The issue: A worker might be uncertain due to poor image quality or lack of expertise. If forced to choose "1" or "0," their internal uncertainty—a valuable data point—is destroyed.
  • Prior Work: Methods like "soft labels" or behavior prediction have tried to fix this, but they often rely on post-hoc guesses about worker reliability rather than capturing uncertainty at the moment of input.

Methodology: The Power of Interval Arithmetic

The core insight of this paper is that an interval contains more "signal" than a single point. If a worker provides a label , they are providing a centroid (the belief) and a radius (the uncertainty).

1. Defining the IVL

An IVL is a sub-interval of .

  • Midpoint (): If , it's a positive lean; if , it's negative.
  • Radius (): Represents the maximum variation or doubt.
  • PDF Generation: Unlike binary labels that follow Bernoulli trials, IVLs allow for the construction of a continuous Probability Density Function (PDF) across all worker responses for a single task.

2. The Inference Algorithm

The authors propose Algorithm 2, which aggregates these intervals into a unified PDF . The final decision is made by calculating: If , the system infers a "1". This ensures that the matching probability is always optimized based on the full distribution of crowd belief.

Model Architecture Placeholder: Finding a PDF for Multiple IVLs The algorithm for calculating the aggregate PDF from a multiset of interval labels.

Experiments and Results

The researchers compared their IVL inference against the industry-standard Majority Voting (MV).

Breaking the Ties

Majority Voting often hits a "tie" when votes are split. IVLs can break these ties by looking at the width and position of the intervals. Even if the count of positive/negative votes is equal, the ranges might lean statistically toward one side.

Resilience to Bias

As shown in the experimental results, when biased labels are introduced, the IVL method effectively shifts the matching probability, providing a quantitative metric for reliability.

Experimental Results: Matching Probability vs Bias Distribution of matching probability with and without bias. Biased labels shift the probability density, allowing for better management of skewed crowd input.

The Uncertainty Index ()

The paper introduces . This index acts as a "health check" for data collection. If is near 0.5, the crowd is completely confused, and the system knows it needs more labels to reach a confident conclusion.

Critical Insight & Conclusion

The true value of this work lies in its Inductive Bias: it assumes that human error is not just random noise to be filtered, but a structured uncertainty to be modeled.

Limitations: The paper primarily uses uniform distributions for intervals. In real-world scenarios, human uncertainty might follow a "Beta" or "Normal" distribution within the interval. Furthermore, the "collusion attack" (where workers coordinate to provide the same biased intervals) remains a challenge that requires separate adversarial detection.

Future Outlook: As we move toward more complex AI alignment tasks (like RLHF), moving away from "A vs B" testing toward "Interval-scale" feedback could significantly reduce the cost of reaching a high-confidence ground truth.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize "soft labeling" or "fuzzy labels" in crowdsourcing to compare their accuracy against interval-valued approaches.
  • What are the foundational theories behind "Interval Arithmetic" in machine learning, and how does this paper's application to label aggregation differ from its use in financial forecasting?
  • Are there any studies exploring the implementation of Interval-Valued Labels in Large Language Model (LLM) Reinforcement Learning from Human Feedback (RLHF) to capture annotator uncertainty?
Contents
Beyond Binary: Managing Crowdsourcing Ambiguity with Interval-Valued Labels
1. TL;DR
2. The Problem: The "Forced Choice" Trap
3. Methodology: The Power of Interval Arithmetic
3.1. 1. Defining the IVL
3.2. 2. The Inference Algorithm
4. Experiments and Results
4.1. Breaking the Ties
4.2. Resilience to Bias
4.3. The Uncertainty Index ($\zeta$)
5. Critical Insight & Conclusion