Crowd IQ: Quantifying Collective Intelligence on Mechanical Turk

Crowd IQ: Measuring the Intelligence of Crowdsourcing Platforms

2012-01-01
Kosinski, M, Bachrach, Y, Kasneci, G, Van-Gael, J, Graepel, T
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces "Crowd IQ," a psychometric framework to measure the intellectual potential of crowdsourcing platforms using Raven’s Standard Progressive Matrices (SPM). By treating Amazon Mechanical Turk (AMT) as a "black box" of collective intelligence, the authors demonstrate that structured crowds can achieve IQ scores as high as 145, outperforming 99.9% of the general population.

TL;DR

Can a group of anonymous workers, paid a few cents per task, be smarter than a PhD? This paper introduces Crowd IQ, a metric derived from Raven’s Standard Progressive Matrices to benchmark the "brainpower" of crowdsourcing platforms. The results are startling: a properly incentivized and aggregated crowd can reach an IQ of 145—surpassing nearly everyone in the general population.

Background: The Invisible Mind

In the early days of crowdsourcing, the primary concern was throughput: How many images can we label per hour? However, as tasks grew more complex—translations, programming, and data analysis—the industry hit a wall. Requesters found that simply throwing money at a problem didn't guarantee quality. The "effort level" of a worker is invisible, leading many to "free-ride" by submitting random or low-effort responses.

This paper shifts the perspective from labor to intelligence. By using a standardized IQ test, the authors provide a universal yardstick to compare humans, crowds, and experimental conditions on the same scale.

The "Crowd IQ" Framework

The methodology involves breaking down the 60-question Raven's SPM test into individual HITs on Amazon Mechanical Turk (AMT). The researchers then manipulated four key levers:

  1. Financial Incentives: Ranging from 0.20 per question.
  2. Reputation Scoring: Filtering workers based on their historical quality.
  3. Social Pressure: Explicitly threatening to reject work (and thus hurt reputation) for incorrect answers.
  4. Aggregation Logic: Using majority vote and dynamic tie-breaking.

Sample HIT Design Figure 1: The task interface presented to AMT workers, emphasizing the "computer-generated reasoning" aspect.

Counter-Intuitive Findings: Why More Money Isn't Better

The study discovered a non-monotone relationship between payment and performance. While increasing pay from 1¢ to 5¢ improved results, further increases to 20¢ actually decreased the Crowd IQ.

The Insight: High rewards attract "free-riders" who gamble for the payout, and may also introduce "choking under pressure," where the high stakes degrade cognitive clarity.

Conversely, the "Threat of Rejection" was the most powerful performance driver. Knowing that a wrong answer could kill their reputation boosted the crowd's IQ from 119 to 138—a massive jump in psychometric terms.

Scaling Intelligence via Aggregation

The paper confirms that "The Wisdom of Crowds" is real, but it requires smart aggregation.

Crowd IQ vs Number of Workers Figure 2: The power of redundancy. As the number of workers per task increases, the Crowd IQ approaches 145.

The researchers went a step further, proposing Adaptive Sourcing. Instead of paying for 5 workers on every question, they only added extra workers when the first two disagreed. This dynamic approach achieved elite-level performance at a fraction of the cost.

Critical Insight & Industry Value

The real takeaway here isn't just that crowds are smart; it's that Crowd IQ is a function of system design.

  • Reputation is Currency: A worker's reputation is more valuable to the requester than the cash reward itself.
  • Targeted Redundancy: Don't treat all tasks as equally difficult. Use adaptive schemes to "spend" your redundancy budget on the hard problems where consensus is elusive.

Limitations & Outlook

While the paper proves that crowds can solve static IQ problems, these workers didn't interact. True "Collective Intelligence" usually involves collaboration (like Wikipedia). Understanding how interactive environments influence Crowd IQ is the next frontier. For modern AI developers, these findings suggest that when building RLHF (Reinforcement Learning from Human Feedback) pipelines, the incentive structure and rejection risk of the human raters are just as important as the model architecture itself.

Conclusion

The crowd is not just a source of cheap labor; it is a scalable cognitive engine. If you design the "market rules" correctly—balancing reputation, moderate pay, and adaptive aggregation—you can tap into a level of intelligence that few individual experts can match.

Find Similar Papers

Try Our Examples

  • Search for recent studies that correlate "Collective Intelligence" factors in remote teams with standardized psychometric testing beyond the Raven's Matrices.
  • How have modern LLM-based evaluation frameworks integrated the "rejection threat" or "reputation mechanisms" originally proposed in early crowdsourcing literature for data labeling?
  • Examine the application of "Adaptive Sourcing" and dynamic tie-breaking algorithms in Large Language Model (LLM) alignment tasks like RLHF.
Contents
Crowd IQ: Quantifying Collective Intelligence on Mechanical Turk
1. TL;DR
2. Background: The Invisible Mind
3. The "Crowd IQ" Framework
4. Counter-Intuitive Findings: Why More Money Isn't Better
5. Scaling Intelligence via Aggregation
6. Critical Insight & Industry Value
6.1. Limitations & Outlook
7. Conclusion