KOG: Building a Structured Wikipedia Ontology via Joint Inference
Automatically refining the wikipedia infobox ontology
KOG (Kylin Ontology Generator) is an autonomous system designed to build a structured ontology by refining Wikipedia infobox classes. It integrates Wikipedia's semi-structured data with WordNet using Markov Logic Networks (MLNs) to achieve SOTA performance in subsumption detection and attribute mapping.
TL;DR
The Kylin Ontology Generator (KOG) transforms the "noisy" collection of Wikipedia infoboxes into a clean, machine-readable ontology. By combining Wikipedia's collaborative edit history with WordNet’s linguistic structure through Markov Logic Networks (MLNs), KOG achieves nearly 99% precision in organizing web-scale knowledge into a formal hierarchy.
Background: The Wild West of Wikipedia Infoboxes
Wikipedia is a goldmine of structured data, with infoboxes providing over 15 million facts. However, this data is essentially a "Wild West":
- Redundancy: Multiple templates like "US County" and "U.S. County" exist for the same concept.
- Inconsistency: One class uses
birth_placewhile another usescity_of_birth. - Lack of Hierarchy: There is no rigorous "is-a" relationship (e.g., knowing that a Particle Physicist is a Physicist).
KOG solves this by treating the infoboxes as classes and the slots as attributes, then applying a sophisticated statistical-relational learning layer to clean and organize them.
Methodology: The Power of Joint Inference
The core technical innovation of KOG is moving beyond simple binary classifiers (like SVMs) toward Joint Inference.
1. Feature Engineering from Heterogeneous Sources
KOG doesn't just look at names. it evaluates:
- Edit History: If users frequently move articles from Class A to Class B, there is likely a taxonomic link.
- Hearst Patterns: Searching the web for phrases like "Scientists such as Physicists" to find hypernyms.
- WordNet Mapping: Linking Wikipedia classes to WordNet nodes to inherit validated linguistic structures.
2. Markov Logic Networks (MLN) Architecture
The beauty of KOG lies in its use of MLNs to resolve the "is-a" hierarchy. Unlike an SVM which looks at pairs of classes in isolation, MLNs allow for global constraints:
- Transitivity: If
A isa BandB isa C, thenA isa C. - Mutual Enhancement: If KOG is confident that
Class Amaps toWordNet Node X, it can use WordNet's hierarchy to better predict Wikipedia's hierarchy.
The architecture illustrates the flow from noisy schema cleaning to the joint inference engine.
Experiments: Why Joint Learning Wins
In head-to-head tests, the MLN approach (labeled MLN+) significantly reduced errors compared to standalone SVMs.
The precision-recall curve shows that the MLN+ model (solid line) maintains higher precision at the same recall levels compared to SVMs.
Key Breakthroughs:
- Cleanup: Successfully renamed obscure templates (e.g., "wfys" to "youth festivals") using spell checks and category tags.
- Attribute Mapping: Achieved 94% precision in mapping attributes across parent-child classes, enabling Faceted Browsing and automated template generation.
Critical Insight: The Value of Semantic Alignment
KOG demonstrates that the true power of Wikipedia isn't just in the facts it contains, but in its ability to be mapped to a formal system like WordNet. By situating "Particle Physicist" under "Physicist," KOG allows a query system to answer questions like "Which scientists won the Nobel Prize?" even if the article only mentions the specific sub-discipline.
Conclusion & Future Outlook
While KOG was a pioneer in using Markov Logic for ontologies, its legacy lives on in how we approach Knowledge Graph Completion today. The primary limitation noted was the sparse data in the "long tail" of Wikipedia, a problem being addressed in modern research by using Large Language Models to "hallucinate" or infer attributes where human editors have been less active.
KOG proved that a "clean" Semantic Web can be built autonomously from the "messy" collaborative efforts of millions of users.
