KOG: Building a Structured Wikipedia Ontology via Joint Inference

Automatically refining the wikipedia infobox ontology

2008-04-21
Fei Wu, Daniel S. Weld
Summary
Problem
Method
Results
Takeaways
Abstract

KOG (Kylin Ontology Generator) is an autonomous system designed to build a structured ontology by refining Wikipedia infobox classes. It integrates Wikipedia's semi-structured data with WordNet using Markov Logic Networks (MLNs) to achieve SOTA performance in subsumption detection and attribute mapping.

TL;DR

The Kylin Ontology Generator (KOG) transforms the "noisy" collection of Wikipedia infoboxes into a clean, machine-readable ontology. By combining Wikipedia's collaborative edit history with WordNet’s linguistic structure through Markov Logic Networks (MLNs), KOG achieves nearly 99% precision in organizing web-scale knowledge into a formal hierarchy.

Background: The Wild West of Wikipedia Infoboxes

Wikipedia is a goldmine of structured data, with infoboxes providing over 15 million facts. However, this data is essentially a "Wild West":

  • Redundancy: Multiple templates like "US County" and "U.S. County" exist for the same concept.
  • Inconsistency: One class uses birth_place while another uses city_of_birth.
  • Lack of Hierarchy: There is no rigorous "is-a" relationship (e.g., knowing that a Particle Physicist is a Physicist).

KOG solves this by treating the infoboxes as classes and the slots as attributes, then applying a sophisticated statistical-relational learning layer to clean and organize them.

Methodology: The Power of Joint Inference

The core technical innovation of KOG is moving beyond simple binary classifiers (like SVMs) toward Joint Inference.

1. Feature Engineering from Heterogeneous Sources

KOG doesn't just look at names. it evaluates:

  • Edit History: If users frequently move articles from Class A to Class B, there is likely a taxonomic link.
  • Hearst Patterns: Searching the web for phrases like "Scientists such as Physicists" to find hypernyms.
  • WordNet Mapping: Linking Wikipedia classes to WordNet nodes to inherit validated linguistic structures.

2. Markov Logic Networks (MLN) Architecture

The beauty of KOG lies in its use of MLNs to resolve the "is-a" hierarchy. Unlike an SVM which looks at pairs of classes in isolation, MLNs allow for global constraints:

  • Transitivity: If A isa B and B isa C, then A isa C.
  • Mutual Enhancement: If KOG is confident that Class A maps to WordNet Node X, it can use WordNet's hierarchy to better predict Wikipedia's hierarchy.

KOG Architecture The architecture illustrates the flow from noisy schema cleaning to the joint inference engine.

Experiments: Why Joint Learning Wins

In head-to-head tests, the MLN approach (labeled MLN+) significantly reduced errors compared to standalone SVMs.

Performance Comparison The precision-recall curve shows that the MLN+ model (solid line) maintains higher precision at the same recall levels compared to SVMs.

Key Breakthroughs:

  • Cleanup: Successfully renamed obscure templates (e.g., "wfys" to "youth festivals") using spell checks and category tags.
  • Attribute Mapping: Achieved 94% precision in mapping attributes across parent-child classes, enabling Faceted Browsing and automated template generation.

Critical Insight: The Value of Semantic Alignment

KOG demonstrates that the true power of Wikipedia isn't just in the facts it contains, but in its ability to be mapped to a formal system like WordNet. By situating "Particle Physicist" under "Physicist," KOG allows a query system to answer questions like "Which scientists won the Nobel Prize?" even if the article only mentions the specific sub-discipline.

Conclusion & Future Outlook

While KOG was a pioneer in using Markov Logic for ontologies, its legacy lives on in how we approach Knowledge Graph Completion today. The primary limitation noted was the sparse data in the "long tail" of Wikipedia, a problem being addressed in modern research by using Large Language Models to "hallucinate" or infer attributes where human editors have been less active.

KOG proved that a "clean" Semantic Web can be built autonomously from the "messy" collaborative efforts of millions of users.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize Markov Logic Networks or Factor Graphs for joint entity and relation extraction in large-scale knowledge bases.
  • What are the current SOTA methods for mapping Wikipedia Infoboxes to the Wikidata or DBpedia ontologies using Deep Learning?
  • How have modern Large Language Models (LLMs) improved the task of 'long-tail' attribute extraction and schema discovery compared to the Kylin system?
Contents
KOG: Building a Structured Wikipedia Ontology via Joint Inference
1. TL;DR
2. Background: The Wild West of Wikipedia Infoboxes
3. Methodology: The Power of Joint Inference
3.1. 1. Feature Engineering from Heterogeneous Sources
3.2. 2. Markov Logic Networks (MLN) Architecture
4. Experiments: Why Joint Learning Wins
4.1. Key Breakthroughs:
5. Critical Insight: The Value of Semantic Alignment
6. Conclusion & Future Outlook