Beyond the Majority: Leveraging Worker Similarity for Expert Discovery in Crowdsourcing
Incorporating Worker Similarity for Label Aggregation in Crowdsourcing
This paper introduces a novel probabilistic label aggregation method for crowdsourcing that incorporates worker similarity to distinguish experts from non-experts. By extending the Generative Model of Labels, Abilities, and Difficulties (GLAD) framework, the authors demonstrate that experts are more likely to reach consensus than non-experts, leading to a state-of-the-art (SOTA) performance on multiple real-world datasets.
TL;DR
In the world of crowdsourcing, we typically assume the majority is right. But what if the "crowd" is mostly noisy non-experts? This paper proposes a new probabilistic approach that identifies experts not just by how often they agree with the majority, but by how similar their labeling patterns are to other potentially high-ability workers. By incorporating squared Jaccard similarity into an EM-based framework, the authors achieve higher accuracy than standard models like GLAD and DARE across diverse real-world datasets.
Problem & Motivation: The "Non-Expert" Trap
The central challenge in crowdsourcing quality control is the cold-start expert problem. Most existing tools (like Dawid-Skene or GLAD) operate on a circular logic: "Correct answers are given by experts, and experts are those who give correct answers." In practice, this often defaults to "experts are those who agree with the majority."
However, the authors point out a critical insight: Experts are consistent in their excellence, while non-experts are random in their errors.
In a 4-choice task, two non-experts have only a 1/4 chance of hitting the same wrong answer, whereas two experts have a near 100% chance of hitting the same correct answer. This "consensus of expertise" provides a signal that is much stronger than mere agreement with the average.
Methodology: Modeling Wisdom through Similarity
The authors extend the classic probabilistic ability model. Instead of treating a worker’s influence as a simple scalar, they define the probability of a worker providing the correct label as:
Where the magic happens in :
eq i} s^2_{ii'} au_{i'}$$ ### Key Components: * **$ au_i$**: The latent ability of worker $i$. * **$s^2_{ii'}$**: The squared Jaccard similarity between worker $i$ and $i'$. Squaring the similarity penalizes low-agreement pairs and rewards high-agreement pairs exponentially. * **Hyperparameter Tuning**: Since ground truth is unknown (unsupervised), the authors use **Perplexity** to find the optimal $\lambda$, effectively measuring how well the model predicts the observed labels.  *Figure 1: Toy example showing how experts (top-left block in similarity matrix c) stand out vividly through consensus.* ## Experiments & Results The authors tested their model against **Majority Voting (MV)**, **GLAD**, and **DARE** across 8 datasets (sentiment analysis, temporal relations, etc.). ### Performance Highlights: * **Expert Identification**: In the *rte* (Recognizing Textual Entailment) dataset, the model achieved **0.9275 accuracy**, outperforming all baselines. * **Resilience to Noise**: In the *duck* dataset, where the average worker accuracy is low, the similarity-weighted model significantly outperformed standard GLAD (0.7685 vs 0.7222).  *Table 2: Comparison of accuracy across different datasets. The similarity-based approach ("Our") wins in the majority of cases.* ## Critical Insight: Quality vs. Quantity A fascinating takeaway from the study of the *duck* and *smile* datasets is that **worker ability trumps redundancy**. The authors found that even with high redundancy (39 workers per item), accuracy remained capped if the average worker ability was low. Their recommendation for practitioners is clear: **Spend your budget on finding and retaining high-ability workers (experts) identified via similarity metrics rather than blindly increasing the number of labels per task.** ## Conclusion & Future Work This work shifts the focus of label aggregation from "who is in the majority" to "who is in the expert cluster." While the current implementation uses Jaccard similarity within an EM framework, the philosophy could easily be extended to Bayesian graphical models (like DARE). **Limitations**: The model struggles when the crowd is exceptionally poor (e.g., the *face* dataset), where even the similarity signal is drowned out by noise. In such cases, some form of basic worker pre-screening remains necessary. --- *Takeaway for the industry: If you're running a labeling campaign, look for the 'Consensus of the Few'. It’s often more accurate than the 'Noise of the Many'.*