Evaluating the Crowd: Insights into Contribution Validation Across 50 Global Platforms

A Review on the Methods to Evaluate Crowd Contributions in Crowdsourcing Applications

2019-11-01
Hazleen Aris, Aqilah Azizan
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a systematic review of 50 crowdsourcing applications to identify and document methods for evaluating crowd contributions. It categorizes initiatives into simple, complex, and creative tasks, identifying "Expert Judgement," "Rating," and "Feedback" as the primary evaluation mechanisms used to ensure data reliability and quality.

TL;DR

Quality control is the "make or break" factor for crowdsourcing. This review of 50 major platforms (from Waze to Innocentive) reveals that while 96% of systems evaluate their users, most still rely on manual human intervention. The study categorizes these methods into Expert Judgement, Rating, and Feedback, providing a roadmap for future automated trust mechanisms.

Contextual Positioning

In the academic coordinate system, this paper serves as a Systematic Taxonomy. It bridges the gap between the theoretical definition of crowdsourcing and the practical implementation of quality assurance (QA). It essentially asks: If we open the doors to everyone, how do we keep the "trash" out?

The Core Problem: The Reliability Gap

Crowdsourcing leverages "human intelligence" for tasks computers struggle with (e.g., assessing washroom cleanliness or solving complex molecular biology problems). However, the "Open Call" nature means:

  • No Guarantees: Contributions may be malicious, lazy, or simply incorrect.
  • Complexity Variance: A method that works for a 5-star restaurant rating doesn't scale to a $1M scientific prize.
  • Manual Bottlenecks: High-quality evaluation currently requires expensive human experts, creating a scalability paradox.

Methodology: Mapping the Crowdsourcing Landscape

The researchers utilized the PRISMA guideline to filter applications. They categorized tasks into three distinct tiers:

  1. Simple Tasks: Distributed human intelligence (e.g., Waze traffic reporting).
  2. Complex Tasks: Specialized problem solving (e.g., Challenge.gov).
  3. Creative Tasks: Idea and design generation (e.g., 99designs).

The Three Pillars of Evaluation

The paper identifies a clear correlation between task type and evaluation method:

  • Expert Judgement (The Gold Standard): Used primarily for Complex and Creative tasks. It is highly reliable but prone to subjective bias and lacks scalability.
  • Rating (The Scalable Choice): Dominates "Simple" tasks. Peer-to-peer ratings (like Airbnb or Blablacar) provide standardized trust but risk "Review Spamming."
  • Feedback (The Context Builder): Qualitative remarks. While useful for iteration (e.g., Crowdspring), it rarely serves as the final filter for selection.

Evaluation Mechanism Distribution

Key Results & Experimental Insights

The analysis of 50 applications produced a striking dataset:

  • Total Coverage: 48 out of 50 apps had formal evaluation.
  • Manual Dominance: Expert judgement appeared in 25/50 apps, often as the sole gatekeeper for high-stakes decisions.
  • The Simple-Rating Link: 10 out of 15 "Simple" apps used Ratings, proving that peer-validation is the industry's go-to for high-volume, low-complexity data.
  • The "Not Evaluated" Risk: Platforms like Kickstarter were noted for lacking internal contribution evaluation, leaving the reliability entirely to the public's "Buyer Beware" intuition.

Application Mapping Table excerpt

Deep Insight: Moving Toward Automation

The most profound takeaway from this review is the Automation Gap.

While high-stakes platforms (Complex/Creative) naturally lean toward Experts, the "Simple" task category is ripe for a technological shift. The authors argue that manual evaluation is:

  1. Error-prone: Human fatigue leads to oversight.
  2. Time-consuming: Dependent on evaluator availability.
  3. Biased: Subjective viewpoints skew results.

Future Industry Impact: The next generation of crowdsourcing will likely integrate Reputation Systems (like the Beta Reputation System) and ML-based anomaly detection to replace manual rating/expert review. This would allow for real-time validation of crowd data, finally solving the scalability paradox.

Conclusion

This review provides the empirical evidence needed to justify the shift toward Automated Quality Assurance. By documenting where the industry stands—heavily reliant on manual expert judgment—it sets the stage for engineers and researchers to build the algorithms that will eventually replace the "Expert" with the "Model."


Index Terms: Evaluation method, Crowdsourcing survey, PRISMA, Human Intelligence.

Find Similar Papers

Try Our Examples

  • Search for recent papers that propose automated evaluation algorithms or machine learning models to validate crowd contributions in simple crowdsourcing tasks.
  • Research the "Beta Reputation System" mentioned in the acknowledgements and how it has been integrated into crowdsourcing trust frameworks since this review.
  • Find comparative studies on the trade-offs between expert-led evaluation and peer-rating systems in terms of cost, bias, and accuracy for creative tasks.
Contents
Evaluating the Crowd: Insights into Contribution Validation Across 50 Global Platforms
1. TL;DR
2. Contextual Positioning
3. The Core Problem: The Reliability Gap
4. Methodology: Mapping the Crowdsourcing Landscape
4.1. The Three Pillars of Evaluation
5. Key Results & Experimental Insights
6. Deep Insight: Moving Toward Automation
7. Conclusion