Double Weighted K-Nearest Voting: Robust Label Aggregation in Crowdsourcing

Double weighted K-nearest voting for label aggregation in crowdsourcing learning

2019-08-30
Jiaye Li, Hao Yu, Leyuan Zhang, Guoqiu Wen
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces the Double Weighted K-Nearest Voting (DWKNV) strategy for label aggregation in crowdsourcing. It combines a worker reliability model with a weighted K-nearest neighbor approach to effectively handle missing labels and low-quality workers, achieving superior accuracy across 12 benchmark datasets.

TL;DR

In crowdsourcing, not all workers are created equal, and not every task gets the same attention. This paper presents a Double Weighted K-Nearest Voting (DWKNV) strategy that simultaneously filters out "bad" workers and fills in the gaps of missing labels by looking at similar tasks. By using a dual-weighting optimization framework, it outperforms traditional Majority Voting and standard KNN methods in both accuracy and robustness.

Problem & Motivation

Crowdsourcing platforms like Amazon Mechanical Turk have revolutionized data labeling, but they suffer from two major flaws:

  1. The "Expert vs. Fraudster" Problem: Traditional Majority Voting (MV) treats a PhD's label the same as a random guesser's.
  2. The Missing Label Problem: If a difficult task receives no labels due to worker avoidance or data loss, MV simply breaks down.

The authors' insight is that similarity is key. If we know Task A is very similar to Task B, we can use the labels from B to help decide the truth for A. Furthermore, by fitting worker responses to the overall sample features, we can statistically "sniff out" workers whose labels don't align with the underlying data structure.

Methodology: The Power of Two Weights

The core of the proposed method lies in a unified objective function that balances worker selection and sample relationships.

1. The Architecture

The model simplifies to a double-voting mechanism:

  • Worker Weight (): A binary or continuous weight that identifies high-performing workers. If a worker's labels don't "fit" the data features well, their weight drops to zero.
  • Neighbor Weight (): Based on the KNN principle, labels from nearby samples are aggregated. The closer the neighbor, the higher its influence on the current sample's final label.

Overall Logic

2. Optimization Strategy

Because the model isn't jointly convex, the authors use an alternate iterative optimization approach:

  • Step 1: Fix the worker weights and optimize the feature coefficient matrix () using Accelerated Proximal Gradient Descent (APGD).
  • Step 2: Fix and update worker weights () by evaluating the fitting error.

This ensures that the algorithm converges to a stable solution where only the most "reliable" labels and workers contribute to the final decision.

Experiments & Results

The authors tested their algorithm on 12 diverse datasets ranging from "Chess" to "German Credit" and "Letter Recognition."

SOTA Comparison

The DWKNV algorithm was compared against:

  • MV: Standard Majority Voting.
  • Knv: Standard K-Nearest neighbor without specific weights.
  • nw-Proposed: Only neighbor weights.
  • ws-Proposed: Only worker weights.

Accuracy Comparison

Key Finding: The "Double Weighted" version systematically beat all other variations. On the CNAE dataset, the proposed method achieved 94.21% accuracy, significantly higher than MV's 90.12%. Even more impressively, in scenarios with missing labels, the KNN component allowed the model to maintain high performance where MV would have failed entirely.

Robustness and Sensitivity

The study also revealed that the model is particularly sensitive to the parameter , which acts as a threshold for worker quality. Adjusting this allows the system to be "strict" (only keeping elite workers) or "lenient" (keeping more workers to maintain volume).

Critical Analysis & Conclusion

The Double Weighted K-Nearest Voting approach is a significant step forward because it moves away from purely frequentist statistics (counting votes) toward a representation-aware aggregation.

Takeaway

By anchoring crowdsourced labels to the actual features of the data samples via KNN, the authors provide a "safety net" for noisy environments. This work proves that the geometry of the data (how samples relate to each other) is just as important as the credibility of the person providing the label.

Limitations & Future Work

While effective, the method relies on having good feature representations () for the samples to calculate distances. If the features are noisy, the KNN weights might become misleading. The authors suggest that future work could incorporate crowdsourcing incentives to improve the initial quality of the labels submitted by workers.


Senior Editor's Note: This paper is a must-read for anyone designing quality control systems for LLM RLHF (Reinforcement Learning from Human Feedback) or large-scale dataset curation.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize graph-based semi-supervised learning for label aggregation in crowdsourcing to compare with the KNN approach.
  • Which paper first proposed the use of fitting learning for worker quality estimation, and how does the sparsity constraint in this paper differ from that origin?
  • Explore how the Double Weighted K-Nearest Voting strategy can be adapted for multi-label crowdsourcing tasks where samples belong to multiple categories simultaneously.
Contents
Double Weighted K-Nearest Voting: Robust Label Aggregation in Crowdsourcing
1. TL;DR
2. Problem & Motivation
3. Methodology: The Power of Two Weights
3.1. 1. The Architecture
3.2. 2. Optimization Strategy
4. Experiments & Results
4.1. SOTA Comparison
4.2. Robustness and Sensitivity
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work