Bridging the Gap: Bootstrapping the Semantic Web via Grammar and Ontology Co-Learning
On the Need to Bootstrap Ontology Learning with Extraction Grammar Learning
The paper proposes a bootstrapping framework that integrates Ontology Learning and Extraction Grammar Learning to automate the creation of Semantic Web resources. By combining the eg-GRIDS genetic algorithm for grammar induction with iterative ontology enrichment, the author aims to overcome the knowledge acquisition bottleneck in Information Extraction (IE).
TL;DR
The vision of the Semantic Web—a machine-readable internet—is stalled by the massive manual effort required to build ontologies and extraction rules. This paper argues for a bootstrapping architecture where Information Extraction (IE) and Ontology Learning exist in a symbiotic loop: ontologies provide the "labels" to train grammar learners, and the resulting grammars discover new "knowledge" to expand the ontologies.
Context & Motivation: The Knowledge Acquisition Bottleneck
Despite decades of research, Information Extraction remains a fragile discipline. Deep NLP methods are often too slow for the Web, while shallow "wrappers" (HTML-based extractors) break the moment a website changes its layout.
The author identifies a fundamental circular dependency:
- To perform Information Extraction, you need a conceptual model (Ontology).
- To build an Ontology at scale, you need Information Extraction to scan the web for concepts.
Instead of trying to solve these as separate problems, this work proposes a self-evolving system managed by Machine Learning.
The CROSSMARC Framework: A Blueprint for Integration
The laboratory (SKEL) developed the CROSSMARC architecture to handle cross-lingual retail data. It serves as the physical implementation of the bootstrapping theory.

The architecture uses independent agents (crawlers, extractors, and coordinators) that communicate via a central blackboard. The Ontology acts as the "language-independent" core that all agents reference, ensuring consistency across English, Greek, French, and Italian.
Methodology: The Engines of Learning
The paper details two specific algorithmic breakthroughs that power this cycle:
1. eg-GRIDS: Genetic Grammar Induction
Inducing Context-Free Grammars (CFGs) from positive examples is notoriously difficult (the "search space explosion" problem). The author utilizes eg-GRIDS, which uses a genetic search strategy guided by the Minimum Description Length (MDL) principle.
- The Intuition: The "best" grammar is the one that most concisely describes the data.
- The Innovation: By replacing beam search with genetic operators, the system searches the space of grammatical patterns an order of magnitude faster, allowing it to handle complex real-world text.
2. Stacked Generalization (Meta-Learning)
Instead of relying on a single "best" extractor, the author uses a stacking framework.

Base-level learners contribute their extractions to a meta-level dataset. A meta-classifier then learns which extraction tool is most reliable for specific entity types (e.g., "Company Name" vs. "Transaction Amount"), effectively breaking the 60% performance barrier common in single-model systems.
The Cycle of Ontology Enrichment
The most compelling part of the methodology is the Ontology Enrichment process.

- Seeding: Start with a small, manually built ontology.
- Auto-Annotation: Use the ontology to find known instances in raw text, labeling them automatically for the IE learner.
- Generalization: The IE system learns patterns (grammars) that describe these instances.
- Discovery: The IE system identifies new candidates that look like concepts but aren't in the ontology yet.
- Expert Review: A human confirms these new concepts, and the cycle restarts with a richer ontology.
Critical Insight & Limitations
This work shifts the focus from "Human-as-Creator" to "Human-as-Validator." However, several challenges remain:
- Knowledge Representation: Should we merge grammars and ontologies into a single hybrid structure (like Conceptual Graphs)?
- Multimedia Integration: How do we extend this bootstrapping logic to non-textual data like images and video?
- Noise Propagation: If the initial IE learner makes a systematic error, it could "pollute" the ontology, creating a negative feedback loop.
Conclusion
Georgios Paliouras makes a strong case that the Semantic Web cannot be "built"—it must be "grown." By leveraging genetic grammar induction and meta-learning within a recursive bootstrapping framework, we can move away from static, labor-intensive systems toward dynamic, self-enriching knowledge bases.
