Corleone: Scaling Entity Matching via Hands-Off Crowdsourcing
Corleone: Hands-off crowdsourcing for entity matching
The paper introduces Corleone, the first Hands-Off Crowdsourcing (HOC) framework for Entity Matching (EM). It automates the entire EM workflow—including blocking, matching, accuracy estimation, and iterative refinement—using an active learning-based Random Forest and crowd workers, eliminating the need for expert developers.
TL;DR
Entity Matching (EM)—the art of identifying different records referring to the same real-world entity—has long been a bottleneck in data science. While crowdsourcing helped, it still required a developer to "babysit" the process. Corleone breaks this barrier by introducing Hands-Off Crowdsourcing (HOC), a system that uses the crowd for every step of the workflow, achieving SOTA accuracy (up to 96.5% F1) without a single line of code from the user.
The "Developer Bottleneck" in Crowdsourcing
Before Corleone, crowdsourced EM followed a "hybrid" model:
- Developer: Writes complex pearl/python scripts for blocking (filtering out obvious non-matches).
- Crowd: Labels a few remaining ambiguous pairs.
- Developer: Evaluates results, tunes the model, and iterates.
For a large enterprise like Amazon or Walmart with thousands of categories, hiring developers for every single EM task is a scaling nightmare. For a journalist or small business owner, it’s an impossible barrier to entry.
Methodology: How Corleone Replaces the Expert
The core innovation of Corleone is its ability to turn "dumb" crowd clicks into "smart" machine-readable rules.
1. Crowdsourced Blocking
How do you get a crowd worker to write a rule like if price_diff > 20 then no_match? You don't. Corleone samples a small portion of the data, uses active learning to build a Random Forest, and then extracts rules from the tree branches. The crowd simply labels pairs, and the system deduces the logic.

2. Active Learning with a "Smart" Stop Button
Training a matcher usually requires thousands of labels. Corleone uses Active Learning to pick the most informative pairs (those where the Random Forest models disagree most).
- The Problem: Crowds are noisy. If you keep training with noisy labels, model accuracy eventually tanks.
- The Solution: A confidence-based stopping mechanism. Corleone monitors a validation set and stops training once the internal "agreement" of the Random Forest peaks, preventing over-exposure to crowd errors.
3. Solving the Skewed Data Problem
In EM, matches are like needles in a haystack (often < 1% of the data). Standard precision/recall estimation fails because a random sample rarely finds enough "needles." Corleone uses its learned rules to "reduce" the haystack, increasing the density of positives so that accuracy can be estimated with 90% fewer labels compared to traditional methods.
Experiments & Results
Corleone was tested against real-world datasets: Restaurants, Citations, and Electronics Products.
- Accuracy: On the "Products" dataset (the hardest one), Corleone achieved an 89.3% F1, smashing the baseline of 69.5%.
- Efficiency: Despite doing everything automatically, the cost was surprisingly low—just 256 for datasets with millions of potential combinations.

Critical Insight: Why This Matters
The takeaway is profound: Crowdsourcing isn't just for labeling data; it's for cleaning and generating models. Corleone proves that an end-to-end HOC system can be "The Godfather" of its domain—managing the complex "mob" of crowd workers in a hands-off fashion to deliver professional-grade results.
Limitations & Future Work
- Crowd Sensitivity: While robust to some noise, a 20% error rate in crowd labels still causes performance drops.
- Cloud Scaling: Future iterations need better Hadoop/Spark integration to handle billions of records in hours rather than days.
Conclusion
Corleone represents a paradigm shift. By removing the developer from the loop, it democratizes high-performance Entity Matching for everyone from Fortune 500 enterprises to "data enthusiasts."
