TripleCheckMate: Bridging the Quality Gap in Linked Open Data through Crowdsourcing
TripleCheckMate: A Tool for Crowdsourcing the Quality Assessment of Linked Data
The paper introduces TripleCheckMate, a specialized tool and methodology for crowdsourcing the quality assessment of Linked Open Data (LOD). It provides a systematic framework for manual evaluation of RDF triples using a predefined 17-criterion taxonomy, successfully identifying and categorizing errors in DBpedia.
TL;DR
The Linked Open Data (LOD) cloud has reached massive proportions, yet the "garbage in, garbage out" principle remains a significant threat. TripleCheckMate is an open-source tool and methodology designed to bring human precision to the wild west of RDF triples. By combining a structured 17-part quality taxonomy with a crowdsourcing workflow, it allows researchers to audit datasets like DBpedia, identify systematic extraction failures, and pave the way for data cleaning.
The Scalability vs. Quality Dilemma
As we push toward 50 billion facts in the LOD world, our reliance on automated extraction (e.g., pulling data from Wikipedia infoboxes) has created a "Quality Gap." While automation provides volume, humans provide context. The authors identify four major dimensions where current LOD often fails:
- Accuracy: Incorrectly extracted object values or datatypes.
- Relevancy: Extracted attributes that are actually layout junk (e.g., image CSS).
- Representational Consistency: Non-standard number or link formats.
- Interlinking: Dead or incorrect links to datasets like Freebase.
Methodology: The Four-Step Audit
The authors don't just provide a tool; they provide a generalized methodology for manual data assessment that can be applied to any knowledge base:
- Selection: Users can focus on specific classes (e.g., "Scientists") or take a random sample to ensure unbiased coverage.
- Evaluation Mode: Determining if the audit is purely manual or assisted by semi-automatic filters.
- Triple Evaluation: This is where TripleCheckMate shines—breaking a resource down into its constituent triples for granular "Right/Wrong" checking.
- Improvement: Translating findings into actual patches via the Patch Request Ontology.
Architecture & Extensibility
TripleCheckMate is built using the Google Web Toolkit (GWT), ensuring a responsive, browser-based experience. Its architecture is intentionally "frontend-heavy" to minimize backend dependencies and maximize portability.

The database schema (as seen below) tracks not just the triple errors, but user sessions and contributor rankings to gamify the process and ensure data lineage.

Real-World Impact: The DBpedia Campaign
In a pilot study, the tool was used to audit DBpedia. The results were revealing:
- 58 users evaluated nearly 3,000 triples.
- The tool supported Inter-rater Agreement (50% chance of double-blind review) to verify if different users identified the same errors.
- Specific DBpedia flaws were identified, such as "Special templates not properly recognized," providing direct feedback to the DBpedia extraction framework developers.

Critical Insight & Future Outlook
The core value of TripleCheckMate isn't just in finding errors, but in categorizing them. By mapping human observations to a formal taxonomy, the authors transform "vague complaints" into "actionable technical requirements."
However, manual crowdsourcing has its limits. The authors acknowledge that the next step is Semi-Automatic Integration. Imagine a system where an AI flags "suspicious" triples based on statistical anomalies, and humans use TripleCheckMate only to verify the most uncertain cases. This hybrid approach will be vital as we move towards even larger Semantic Web structures in the AI era.
Conclusion
TripleCheckMate proves that for high-stakes data—like scientific or cultural knowledge—the human element remains irreplaceable. By providing an extensible, open-source framework, the authors have given the LOD community a fundamental tool for turning a "quantity of data" into a "quality of knowledge."
