CrowdDB: Bridging the Gap Between SQL and Human Intelligence
CrowdDB: Answering Queries with Crowdsourcing
CrowdDB is a pioneering hybrid relational database system that integrates crowdsourcing into the query processing engine to handle tasks machines cannot solve. It extends SQL into "CrowdSQL," utilizing human intelligence via platforms like Amazon Mechanical Turk for data discovery, entity resolution, and subjective ranking.
TL;DR
CrowdDB is a hybrid database system that breaks the "Closed World Assumption" of traditional RDBMS. By extending SQL with crowdsourcing primitives, it allows queries to span beyond stored data to find missing information, resolve ambiguous entities, and perform subjective rankings using human workers from platforms like Amazon Mechanical Turk.
The "Literal" Wall: Why Standard Databases Fail
Traditional Relational Database Management Systems (RDBMS) are incredibly efficient but fundamentally "literal." They suffer from two main limitations:
- The Closed World Assumption: If a record isn't in the table, it simply doesn't exist. A query for "IBM's market cap" returns empty if the record was entered as "International Business Machines."
- Lack of Subjective Intuition: Machines cannot natively answer which image "best visualizes Business Success" without pre-existing labels.
CrowdDB acknowledges that while machines excel at bulk data processing, humans excel at Data Discovery (finding things via search) and Data Comparison (fuzzy matching and ranking).
Methodology: Engineering the Crowd as a Processor
The core innovation of CrowdDB is treating the crowd not as a separate service, but as a specialized query operator within the database engine.
1. CrowdSQL: Extending the Language
The authors introduced simple yet powerful DDL/DML extensions:
CROWDColumns/Tables: Marking an attribute asCROWDtells the engine that if a value is missing (CNULL), it should be fetched from a human.- Subjective Operators:
CROWDEQUAL(~=) for entity resolution andCROWDORDERfor rankings based on human perception.
2. The Architecture
CrowdDB integrates a HIT Manager and a User Interface Manager directly into the query pipeline.

When a query requires human input, the system:
- Generates an HTML interface based on the table schema.
- Batches multiple "jobs" into a Human Intelligence Task (HIT).
- Posts the task to Amazon Mechanical Turk (AMT).
- Collects results and applies majority voting to ensure quality.
Experimental Insights: Performance vs. Reward
The researchers conducted over 25,000 HITs to understand the "Human Processor." Key findings include:
- Throughput vs. Latency: Larger HIT Groups (more tasks posted at once) significantly reduce the time to get the first result because they are move visible to workers. However, group sizes of 50-100 provided the best overall completion rate.
- The Price of Speed: Increasing the reward from 1 cent to 4 cents per task drastically improved completion speed, but paying 2 cents vs 3 cents showed negligible difference, suggesting a non-linear relationship between incentive and performance.

Real-World Utility: Beyond Logic
CrowdDB was tested on subjective tasks like ordering pictures of the "Golden Gate Bridge." The results closely matched expert rankings, proving that the system can capture human aesthetics and "common sense" that machines lack.

Critical Insight & Limitations
While CrowdDB successfully hides the complexity of crowdsourcing from the developer (offering Physical Data Independence for the crowd), it introduces new challenges:
- Open World Uncertainty: Since the crowd can provide an infinite number of new tuples, queries can potentially run forever or exceed budgets if
LIMITclauses aren't strictly used. - Worker Affinity: Humans are not "fungible" like CPUs. They have memories, feelings, and reputation systems (like Turker Nation). Rejecting too many tasks for quality control can lead to a requester being blacklisted by the community.
Conclusion
CrowdDB represents a significant step toward a truly "intelligent" database. By blending the strict logic of SQL with the heuristic flexibility of human workers, it enables a new class of applications that were previously impossible to automate. Future work will likely focus on cost-based optimization—treating "human time" and "dollars" as variables in the standard query optimization equation.
