Artequakt: Bridging the Semantic Gap with Ontology-Driven Knowledge Extraction

Natural Language Processing

2026-01-01
Salvatore Claudio Fanni, Domitilla Deri, Francesca Pia Caputo, Alessio Guarracino, Ilaria Ambrosini, Emanuele Neri, Dania Cioni
Summary
Problem
Method
Results
Takeaways

This paper introduces Artequakt, a system for the automatic extraction of domain-specific knowledge from unstructured Web documents guided by an ontology. By combining IE tools like GATE with WordNet-based lexical expansion, it populates a knowledge base to generate personalized, narrative biographies of artists.

TL;DR

Artequakt is a pioneering system that automates the transition from unstructured Web text to a structured Semantic Web. By using an ontology as a guide, it extracts not just names but meaningful relationships from the lives of artists, storing them in a Knowledge Base (KB) to generate personalized narratives. It solves the scalability issue of manual annotation by "teaching" the extraction tool what to look for through domain-specific schemas.

Problem & Motivation: The Annotation Bottleneck

The vision of the Semantic Web requires rich, machine-readable metadata. However, we face a "chicken and egg" problem: advanced knowledge services need annotations, but manual annotation is too expensive to scale across the billions of pages on the Web.

Existing Information Extraction (IE) tools are excellent at recognizing "named entities" (e.g., identifying that "Rembrandt" is a person), but they are notoriously bad at understanding relationships (e.g., "Rembrandt was born in Leiden"). Without these relations, the data is a collection of isolated facts rather than a coherent web of knowledge. The authors argue that we need an ontology-driven approach where the schema of the domain itself guides the extraction process.

Methodology: The Architecture of Extraction

Artequakt's workflow is divided into three core stages: Extraction, Storage, and Narrative Generation.

1. The Strategy of Guided Extraction

Instead of using static templates, Artequakt uses its ontology (modeled after the CIDOC CRM) to query for what it expects to find. To handle the linguistic variety of the Web, it employs WordNet. If the ontology looks for the relation depict, the system uses WordNet to also search for synonyms like portray or hypernyms like represent.

The Artequakt architecture

2. Syntactic and Semantic Analysis

The system pipeline follows a sophisticated NLP path:

  • Filtering: Using vector similarity to discard irrelevant sites (e.g., filtering out "Rembrandt Hotels" when looking for the painter).
  • Parsing: Utilizing the Apple Pie Parser to group grammatically related phrases.
  • Triplet Generation: Mapping parsed sentences into (Subject, Relation, Object) triples that match the ontology's class structure.

Knowledge Extraction Example

Experiments & Results: From Triples to Tales

The authors tested Artequakt on a corpus of 100 Web pages focused on Impressionist artists. The extraction tool did more than just find data; it consolidated it. By identifying duplicates and resolving identities, it distilled thousands of raw mentions into a clean KB of 600 unique relations.

The "holy grail" of this system is the Narrative Generation. Using "Biography Templates," the system queries the KB to build a story. If a high-quality paragraph about an artist's birth is found, it is inserted directly. If only raw facts exist (e.g., "Born: 1606"), the system dynamically generates a sentence.

Final rendered biography

Critical Analysis & Conclusion

Takeaway

Artequakt represents a significant shift from "passive" storage to "active" acquisition. It proves that an ontology can act as a lens, focusing NLP tools on the information that truly matters for a specific domain.

Limitations

While powerful for the early 2000s, the system faces challenges with referential integrity—if two different sources describe the same event in conflicting ways, the system must decide which to trust. Furthermore, the reliance on WordNet expansion, while clever, can still struggle with highly metaphorical or context-heavy language.

Future Outlook

Today, this work serves as an ancestor to modern Retrieval-Augmented Generation (RAG) and Knowledge Graph (KG) integration. The core insight—that narratives should be constructed from a verified "world model" (the ontology) rather than just pulled from raw text—remains a cornerstone of trustworthy AI research.

Find Similar Papers

Try Our Examples

  • Find recent papers that extend ontology-based knowledge extraction using Large Language Models (LLMs) instead of traditional NLP parsers like GATE.
  • What are the state-of-the-art methods for "entity consolidation" and "referential integrity" when merging knowledge from multiple overlapping Web sources?
  • Explore how narrative generation techniques from structured knowledge bases have evolved since the use of fixed templates like Auld Linky.
Contents
Artequakt: Bridging the Semantic Gap with Ontology-Driven Knowledge Extraction
1. TL;DR
2. Problem & Motivation: The Annotation Bottleneck
3. Methodology: The Architecture of Extraction
3.1. 1. The Strategy of Guided Extraction
3.2. 2. Syntactic and Semantic Analysis
4. Experiments & Results: From Triples to Tales
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook