Competitive Crowdsourcing: How Ranking and Competition Can Slash Labeling Costs by 60%

Competitive Game Designs for Improving the Cost Effectiveness of Crowdsourcing

2014-11-03
Markus Rokicki, Sergiu Chelaru, Sergej Zerr, Stefan Siersdorfer
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces and evaluates competitive and randomized reward mechanisms for crowdsourcing tasks, moving beyond the traditional "Pay-per-HIT" model. By implementing "Winner-Takes-It-All" and "Exponential Reward" schemes, the authors achieved a 3x increase in annotation throughput and a 2.5x reduction in cost per unit for captcha translation.

TL;DR

Researchers from the L3S Research Center have proved that the standard way we pay crowd workers—per task—is fundamentally inefficient. By introducing Exponential Reward Schemes and Medium Information Policies (where you only see people slightly better than you), they increased annotation volume by 300% and significantly lowered the cost per label in both simple (Captcha) and complex (Face Recognition) tasks.

Background: Beyond the Linear Paycheck

In the world of Amazon Mechanical Turk (MTurk), the status quo is simple: do one HIT, get . While reliable, this "Pay-per-HIT" model lacks the psychological "hook" that drives people to excel. This paper asks a provocative question: Can we borrow mechanics from lotteries, eSports, and the Netflix Prize to get more value for the same budget?

The Core Insight: The "Medium" Information Sweet Spot

One of the most profound contributions of this work is the analysis of Information Policies.

  • Open Policy: Seeing the #1 person with 50,000 points when you have 10 is demoralizing.
  • Restricted Policy: Working in a vacuum eliminates the competitive urge.
  • Medium Policy: By showing workers their rank and their immediate neighbors in the leaderboard, the authors triggered a "catch-up" or "defense" instinct that kept workers engaged longer and at higher speeds.

Methodology: Gaming the Reward System

The authors tested several structures against a standard baseline:

  1. Winner-Takes-It-All (WTA): The top performer gets the entire $50 budget.
  2. Exponential (EXP): Rewards are spread across the top 10, but the #1 spot gets significantly more than #10 (e.g., 0.50).
  3. Lottery (Lott): Workers earn tickets based on work, and winners are drawn randomly.

Architectural Flow

Contribution of Individual Workers Figure 1: Notice how Exponential strategies (EXP) result in a more "balanced" contribution from the top 10 compared to Winner-Takes-It-All, where lower ranks give up sooner.

Experimental Showdown

1. Captcha Translation (The Speed Test)

This was a high-volume, "dull" task designed to measure pure monetary motivation. The results were staggering. The exp-med (Exponential reward + Medium info) configuration yielded 150,000+ captchas for a $50 budget, while the baseline only achieved ~58,000.

Table 1: Performance Comparison Table 1: The 'exp-med' strategy reduced cost per captcha to 0.032 cents, compared to 0.077 cents in the baseline.

2. Face Recognition (The Quality Test)

Does speed kill quality? To test this, the authors used the PubFig database. They used "honeypots" (test questions with known answers) to penalize bad actors.

  • Baseline: 660 correct matches per USD.
  • Competitive (Exp-Med): 930 correct matches per USD.

The accuracy remained consistent (fixed at ~92%), proving that competition increases efficiency without necessarily inducing "rushed" errors.

Temporal Dynamics: The "Crossing Curves"

The paper highlights a fascinating psychological phenomenon in the competitive groups. By tracking worker activity over time, they observed "crossing curves."

Worker Activity Curves Figure 5d: In the exp-med strategy, workers constantly leapfrog each other's ranks. This indicates a "fierce" and active competition that persists throughout the duration of the experiment.

Critical Perspective & Limitations

While the gains are undeniable, the paper touches on the ethical implications of "gamifying" labor. Competitive designs can lead to "worker fatigue" and stress, especially if workers depend on these platforms for their living wage. Additionally, the lottery-based systems performed poorly in this study, likely because the prizes weren't high enough to trigger "risk-seeking" behavior.

Conclusion

This research provides a blueprint for any AI researcher or data manager looking to scale labeling efforts. If you want results:

  1. Don't pay linearly. Use a convex, exponential reward curve for top performers.
  2. Provide "Local" Feedback. Don't just show a global leaderboard; show workers who they are about to overtake.
  3. Use Honeypots. Competition drives speed; automated quality checks are required to ensure that speed doesn't compromise the integrity of your training data.

Find Similar Papers

Try Our Examples

  • Find recent studies on "gamified crowdsourcing" that specifically compare individual vs. team-based competitive reward structures.
  • Which theoretical papers first established the "principal-agent" model in crowdsourcing, and how does this paper's empirical ranking-based approach deviate from those economic foundations?
  • Search for research investigating the long-term effects of "competitive reward fatigue" on crowd worker retention and data quality over several months.
Contents
Competitive Crowdsourcing: How Ranking and Competition Can Slash Labeling Costs by 60%
1. TL;DR
2. Background: Beyond the Linear Paycheck
3. The Core Insight: The "Medium" Information Sweet Spot
4. Methodology: Gaming the Reward System
4.1. Architectural Flow
5. Experimental Showdown
5.1. 1. Captcha Translation (The Speed Test)
5.2. 2. Face Recognition (The Quality Test)
6. Temporal Dynamics: The "Crossing Curves"
7. Critical Perspective & Limitations
8. Conclusion