Bridging the Semantic Gap: Ontology-Driven Multi-Relational Data Mining
Ontology Model for Multi-relational Data Mining Application
The paper proposes an ontology-based framework for Multi-Relational Data Mining (MRDM) using the OWL (Web Ontology Language) and the Protégé editor. It demonstrates how formalizing domain knowledge through a meta-model and sub-models can improve the relevance of discovered patterns in complex clinical databases, specifically applied to a large hospital in Curitiba.
TL;DR
Standard data mining often struggles with the sheer complexity of multi-table relational databases. This paper introduces a structured methodology using OWL (Web Ontology Language) to build a bridge between raw data and business logic. By defining a global meta-model and specific sub-models, the authors demonstrate how a formal ontology can guide the data mining process, making it more efficient and semantically relevant, specifically within the complex environment of a hospital's clinical records.
Problem & Motivation: The Context Crisis
In the early 2000s, the field of Multi-Relational Data Mining (MRDM) emerged to move beyond the limitations of single-table analysis. However, as the authors point out through their review of ACM KDD workshops (2002-2007), most solutions were purely algorithmic—relying on Inductive Logic Programming (ILP) or probabilistic models.
The unresolved pain point was the lack of "explicit context." Without a formal conceptual model, algorithms spend excessive computational effort exploring irrelevant branches of the database schema. The semantics remained "hidden" within the code, making the results difficult for domain experts (like doctors or nurses) to interpret or validate.
Methodology: From Tables to Ontologies
The core of this research is a systematic 10-step process designed to extract meaning from relational metadata and transform it into a formal ontology.
The 10-Step Reengineering Process
- Metadata Identification: Locating PKs, FKs, and bridge tables.
- Bond Detection: Mapping relationships based on foreign keys.
- Attribute Equivalence: Checking data types and names across tables.
- Value Analysis: Analyzing actual data values to detect hidden equivalences.
- Categorization: Sorting data into Nominal and Numerical types.
- Cleaning: Removing null-heavy columns.
- Temporal Splitting: Breaking Dates/Times into granular components (Day, Month, Year).
- Grouping: Consolidating attributes with similar characteristics.
- Summarization: Creating statistical profiles (Mean, Min, Max) for numerical data.
- Synthetic Meta-Model: Generating a unified view for business analysis.
Tooling & Technology
The authors leveraged a professional stack including ErWin for database reengineering, DbDesigner for sub-model creation, and Protégé for developing the OWL-based ontology.
Figure 1. The relationship between the Global Meta-model and Sub-models.
Real-World Application: The Red Cross Hospital Case Study
The methodology was tested on the database of the Red Cross Hospital in Curitiba. The "Monolithic" database was broken down into three logical pillars:
- Patient Care: Tracking admissions and demographics.
- Prescription: Modeling the medical logic of drug administration.
- Solution: Tracking the outcomes and clinical responses.
By using OWL, the authors created a "network of knowledge." Instead of the mining algorithm blindly traversing tables, it followed the pathways defined by the ontology.
Figure 2. A detailed view of the Patient Care sub-model, showing the granularity utilized for mining.
Critical Insight: The Global Functional Schema (EFG)
The paper introduces the Global Functional Schema (EFG) as the "common denominator" in a multi-relational environment. This schema acts as a translation layer. It allows for the representation of knowledge with varying degrees of precision or uncertainty, which is a common reality in medical data.
Conclusion & Future Outlook
The integration of ontologies into data mining transforms the process from a "black box" search into a guided exploration. By formalizing the domain knowledge before the mining begins:
- Search efficiency is increased (less branching).
- Human-interpretability is improved (aligned with business units).
- Data quality is ensured through systematic reengineering.
Limitations: While powerful, building an OWL ontology requires significant domain expertise and manual effort up-front. Future research might look into "Automated Ontology Learning" to speed up the 10-step process described here.
