CrowdEC: Reducing Costs and Redundancy in Crowdsourced Entity Collection

Incentive-Based Entity Collection Using Crowdsourcing

2018-04-01
Chengliang Chai, Ju Fan, Guoliang Li
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces CrowdEC, an incentive-based crowdsourced entity collection framework designed to gather complete and correct entity sets (e.g., NBA players, universities) while minimizing costs. It achieves state-of-the-art performance by combining a worker elimination strategy with a dynamic bonus-based incentive pricing mechanism.

Executive Summary

CrowdEC is a sophisticated framework designed to solve the "popularity bias" in crowdsourcing, where workers tend to provide the same easy-to-recall entities (e.g., Lebron James) while ignoring the long-tail (e.g., bench players). By integrating a Worker Elimination module and an Incentive Pricing mechanism, the framework ensures that we only pay for high-quality, distinct data. This work represents a significant leap from passive statistical estimation to active participant management in the crowdsourcing ecosystem.

The Pain Point: The Expensive Duplication Problem

Crowdsourcing is the go-to method for building knowledge bases. However, it faces two structural flaws:

  1. Wasteful Redundancy: If you ask 100 workers for an NBA player, 90 might say "Stephen Curry." You've paid for 100 tasks but gained almost no new knowledge.
  2. The Quality Vacuum: In open-ended tasks ("Give me a university name"), there is no pre-defined "Golden Set" to test worker honesty, leading to "cheating" or low-effort submissions.

Methodology: The Architecture of Incentives

CrowdEC shifts the paradigm by treating entity collection as a dynamic optimization problem.

1. Worker Elimination (Utility Maximization)

Not all workers are equal. CrowdEC calculates a Worker Utility score based on:

  • Throughput ( / ): How many unique entities a worker provides per request.
  • Error Bounding (): Using a Bayesian approach to estimate the probability that a submitted entity actually belongs to the target domain ().

CrowdEC Framework Architecture

2. Incentive Pricing (The Bonus Strategy)

Since platforms like Amazon Mechanical Turk (AMT) don't support dynamic price negotiation, CrowdEC uses a Bonus-based retry logic:

  • NoBonus: Standard task, base reward.
  • Bonus: If the worker provides a duplicate, the system notifies them and offers a bonus for a distinct new entry. This turns a "failed" attempt into a successful data point at a lower marginal cost than a new task.

Pricing Interfaces

Experiments and Results

The authors validated the system using datasets like ActiveNBA (450 entities) and TopUniv (100 entities).

  • Cost Efficiency: CrowdEC significantly flattened the cost curve. By the time 420 NBA players were collected, CrowdEC's costs were nearly 3x lower than traditional "Enumeration" methods.
  • Precision/Recall Balance: Unlike "No-Selection" baselines where precision drops as workers get tired or bored, CrowdEC’s quality control kept precision above 95%.

Cost Comparison Results

Critical Analysis & Takeaways

The brilliance of CrowdEC lies in its Online Algorithm. It makes pricing decisions in real-time without needing to know the future distribution of worker requests, achieving an approximation ratio close to the theoretical optimum (Theorem 2).

Future Outlook: While highly effective, CrowdEC assumes that "entity resolution" (deciding if "S. Curry" and "Stephen Curry" are the same) is handled externally. Integrating a real-time, human-in-the-loop entity resolution module could further refine the "Distinctness" metric and provide even higher cost savings for massive datasets like the 4,260-item "AllNBA" list.

Conclusion: CrowdEC proves that financial incentives, when applied algorithmically based on data distinctness rather than just participation, can solve one of the oldest problems in human-powered data collection.

Find Similar Papers

Try Our Examples

  • Search for recent studies on "open-world entity collection" that use reinforcement learning to optimize worker incentive strategies.
  • Which paper originally defined the "Chao92" species estimation model, and how does CrowdEC adapt this formula to handle incorrect human inputs?
  • Investigate how incentive pricing models like those in CrowdEC are applied to multi-modal crowdsourcing tasks, such as large-scale image labeling or audio transcription.
Contents
CrowdEC: Reducing Costs and Redundancy in Crowdsourced Entity Collection
1. Executive Summary
2. The Pain Point: The Expensive Duplication Problem
3. Methodology: The Architecture of Incentives
3.1. 1. Worker Elimination (Utility Maximization)
3.2. 2. Incentive Pricing (The Bonus Strategy)
4. Experiments and Results
5. Critical Analysis & Takeaways