CrowdMD: Scaling Data Veracity via Hybrid Rule Generation
CrowdMD: Crowdsourcing-based approach for deduplication
CrowdMD is a hybrid machine-crowdsourcing framework designed to generate Matching Dependencies (MDs) for data deduplication and entity resolution. By combining human intelligence for semantic labeling with an Apriori-inspired algorithm for rule mining, it achieves high-quality data cleaning rules while minimizing expensive manual expert labor.
TL;DR
CrowdMD introduces a clever hybrid pipeline that tackles the "Veracity" problem in Big Data by generating Matching Dependencies (MDs)—rules that dictate when two records represent the same entity. Instead of asking expensive experts or overloading crowdsourcers with millions of pairs, it uses a small "human-labeled" sample to seed a machine-learning algorithm that mines robust deduplication rules.
Problem & Motivation: The Expert Bottleneck
In the era of Big Data, data quality (De-duplication or Entity Resolution) is traditionally a manual nightmare. Previous SOTA methods followed two paths:
- Expert-Led: Domain experts manually define rules. This is accurate but fails when data velocity is high.
- Crowd-Only: Every pair is sent to the crowd. This is bankruptingly expensive at scale.
The authors of CrowdMD realized that the crowd shouldn't just be "detectors"—they should be "teachers." By labeling a small sample, the crowd can teach the machine how to build rules that apply to the rest of the billion-row database.
Methodology: The Hybrid Engine
CrowdMD splits the labor based on what humans and machines do best.
1. Human Intuition (The RHS)
Humans are naturally good at seeing "The Chipiron" and "Le Chipiron" and knowing they are the same restaurant. CrowdMD asks workers to identify these pairs and—crucially—point out which columns (like Address or Spec) are inconsistent and should be merged. This forms the Right-Hand Side (RHS) of the rule.
2. Algorithmic Mining (The LHS)
Once humans confirm which pairs are duplicates, the machine takes over to find the Left-Hand Side (LHS). It builds a Similarity Table (calculating metrics like Jaro-Winkler or Levenshtein distance) and runs a modified Apriori Algorithm.

The algorithm identifies the minimal subset of attributes that, when similar, reliably lead to the match identified by the crowd. This avoids "over-fitting" a rule to a single pair.
Quantitative Strategy: Managing Crowd Error
A major challenge in crowdsourcing is "noisy" workers. CrowdMD doesn't treat all votes equally. It calculates a probability for each match based on a weighted sum of worker error rates ():
- Error Rate Formula: (Ratio of correct to wrong previous answers).
- Decision Logic: A pair is only marked as a duplicate if the weighted probability of the "Match" votes outweighs the "No Match" votes.
Experiments & Results
In a demonstration using a Restaurant database, the system showed it could generate complex rules like:
Restaurants[Name] ≈ Restaurants[Name] ∧ Restaurants[Tel] ≈ Restaurants[Tel] → Restaurants[Spec, Add] ⇀↽ Restaurants[Spec, Add]
This rule means: "If the Name and Phone are similar, merge the Speciality and Address." This allows the machine to automate millions of matches that were previously impossible due to slight typos in addresses or missing specialty info.

Deep Insight & Conclusion
The real contribution of CrowdMD isn't just "using the crowd," but using a Rule-Based Framework as an intermediate. Most entity resolution systems provide a "black box" list of matches. By generating MDs, CrowdMD provides a transparent, auditable, and reusable asset (the ruleset) that can be applied to new data without bothering the crowd again.
Limitations: The current approach relies on a similarity threshold (fixed at 0.7 in the example). Defining these thresholds automatically remains a challenge for future iterations.
Future Outlook: As data grows, the "Data Reconciliation Rules" proposed by the authors could become the standard for Linked Open Data (LOD) and automated data lake curation.
