CrowdMD: Scaling Data Veracity via Hybrid Rule Generation

CrowdMD: Crowdsourcing-based approach for deduplication

2015-10-01
Asma Abboura, Soror Sahri, Mourad Ouziri, Salima Benbernou
Summary
Problem
Method
Results
Takeaways
Abstract

CrowdMD is a hybrid machine-crowdsourcing framework designed to generate Matching Dependencies (MDs) for data deduplication and entity resolution. By combining human intelligence for semantic labeling with an Apriori-inspired algorithm for rule mining, it achieves high-quality data cleaning rules while minimizing expensive manual expert labor.

TL;DR

CrowdMD introduces a clever hybrid pipeline that tackles the "Veracity" problem in Big Data by generating Matching Dependencies (MDs)—rules that dictate when two records represent the same entity. Instead of asking expensive experts or overloading crowdsourcers with millions of pairs, it uses a small "human-labeled" sample to seed a machine-learning algorithm that mines robust deduplication rules.

Problem & Motivation: The Expert Bottleneck

In the era of Big Data, data quality (De-duplication or Entity Resolution) is traditionally a manual nightmare. Previous SOTA methods followed two paths:

  1. Expert-Led: Domain experts manually define rules. This is accurate but fails when data velocity is high.
  2. Crowd-Only: Every pair is sent to the crowd. This is bankruptingly expensive at scale.

The authors of CrowdMD realized that the crowd shouldn't just be "detectors"—they should be "teachers." By labeling a small sample, the crowd can teach the machine how to build rules that apply to the rest of the billion-row database.

Methodology: The Hybrid Engine

CrowdMD splits the labor based on what humans and machines do best.

1. Human Intuition (The RHS)

Humans are naturally good at seeing "The Chipiron" and "Le Chipiron" and knowing they are the same restaurant. CrowdMD asks workers to identify these pairs and—crucially—point out which columns (like Address or Spec) are inconsistent and should be merged. This forms the Right-Hand Side (RHS) of the rule.

2. Algorithmic Mining (The LHS)

Once humans confirm which pairs are duplicates, the machine takes over to find the Left-Hand Side (LHS). It builds a Similarity Table (calculating metrics like Jaro-Winkler or Levenshtein distance) and runs a modified Apriori Algorithm.

CrowdMD Architecture

The algorithm identifies the minimal subset of attributes that, when similar, reliably lead to the match identified by the crowd. This avoids "over-fitting" a rule to a single pair.

Quantitative Strategy: Managing Crowd Error

A major challenge in crowdsourcing is "noisy" workers. CrowdMD doesn't treat all votes equally. It calculates a probability for each match based on a weighted sum of worker error rates ():

  • Error Rate Formula: (Ratio of correct to wrong previous answers).
  • Decision Logic: A pair is only marked as a duplicate if the weighted probability of the "Match" votes outweighs the "No Match" votes.

Experiments & Results

In a demonstration using a Restaurant database, the system showed it could generate complex rules like: Restaurants[Name] ≈ Restaurants[Name] ∧ Restaurants[Tel] ≈ Restaurants[Tel] → Restaurants[Spec, Add] ⇀↽ Restaurants[Spec, Add]

This rule means: "If the Name and Phone are similar, merge the Speciality and Address." This allows the machine to automate millions of matches that were previously impossible due to slight typos in addresses or missing specialty info.

Crowd Interface Example

Deep Insight & Conclusion

The real contribution of CrowdMD isn't just "using the crowd," but using a Rule-Based Framework as an intermediate. Most entity resolution systems provide a "black box" list of matches. By generating MDs, CrowdMD provides a transparent, auditable, and reusable asset (the ruleset) that can be applied to new data without bothering the crowd again.

Limitations: The current approach relies on a similarity threshold (fixed at 0.7 in the example). Defining these thresholds automatically remains a challenge for future iterations.

Future Outlook: As data grows, the "Data Reconciliation Rules" proposed by the authors could become the standard for Linked Open Data (LOD) and automated data lake curation.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize active learning to further reduce the number of crowdsourcing iterations in Matching Dependency discovery.
  • Which original papers defined the formal semantics of Matching Dependencies (MDs), and how does current research extend these to probabilistic MDs?
  • Explore studies that apply hybrid crowdsourcing-algorithmic entity resolution techniques specifically to multi-modal or unstructured Big Data sources.
Contents
CrowdMD: Scaling Data Veracity via Hybrid Rule Generation
1. TL;DR
2. Problem & Motivation: The Expert Bottleneck
3. Methodology: The Hybrid Engine
3.1. 1. Human Intuition (The RHS)
3.2. 2. Algorithmic Mining (The LHS)
4. Quantitative Strategy: Managing Crowd Error
5. Experiments & Results
6. Deep Insight & Conclusion