DoMO: Democratizing Data Science through Multi-Level Ontologies

Domain-Oriented Multilevel Ontology for Adaptive Data Processing

2020-01-01
Tianxing Man, Elena N. Stankova, Alexander Vodyaho, Nataly Zhukova, Yulia A. Shichkina
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces DoMO (Domain-oriented Multi-level Ontology), a comprehensive semantic framework designed to bridge the gap between complex data mining algorithms and non-expert researchers. By integrating existing ontologies like DMOP and OntoDT into a four-level architecture, it enables automated, adaptive selection of data processing workflows and algorithms based on specific dataset characteristics.

TL;DR

The complexity of modern data mining (DM) often leaves non-expert researchers lost in a sea of algorithms. DoMO (Domain-oriented Multi-level Ontology) is a breakthrough framework that uses semantic meta-mining to automatically recommend the best data processing workflows. By mapping dataset metadata to algorithm strengths, it transforms data science from a "trial-and-error" black box into a logical, queryable knowledge base.

Background: The Gap in Intelligent Data Assistants

The "No Free Lunch" theorem dictates that no single algorithm is best for every task. In the world of Data Mining (DM), selecting the right workflow (Preprocessing -> Algorithm -> Evaluation) requires deep expertise. While Meta-learning aims to solve this by learning from past experiments, it often treats algorithms as black boxes.

Semantic Meta-mining offers a superior alternative: it uses ontologies to understand the internal mechanisms of algorithms and the intrinsic properties of data. However, previous ontologies like DMOP or OntoDM were either too narrow in scope or too complex for practical use.

Methodology: The DoMO Architecture

DoMO's power lies in its multi-level structure, which organizes knowledge from abstract data types to specific implementations.

DoMO Architecture

1. The Four Layers of Insight

  • Upper Level: Uses OntoDT to define the fundamental DNA of data (basic types, value spaces).
  • Domain Level: Where experts define what "Short," "Long," or "Noisy" means for their specific field (e.g., Time Series vs. Image Recognition).
  • Application Level: A specific core ontology tailored for a domain.
  • Implementation Level: The "User Interface" where natural queries are translated into SPARQL-like logic to find executable code.

2. The INPUT Ontology: The User Bridge

The researchers created a specialized INPUT Ontology. This acts as a semantic interface where a user can describe their dataset in plain terms (e.g., "Dataset with 1000 samples, high noise, 3 classes") and receive a full-stack algorithm recommendation.

Case Study: Time Series Classification (TSC)

To prove DoMO's effectiveness, the authors applied it to the CinCECGtorso dataset (ECG data).

The Workflow:

  1. Describe: The user identifies the dataset as SmallTrainTSDataset, LongTSDataset, and ECGTSDataset.
  2. Query: The system searches for algorithms suitableFor these specific constraints.
  3. Recommend: It identified BOSS, COTE, EE, MSM_1NN, and ST as optimal candidates.

Experimental Evidence

When the authors benchmarked all available TSC algorithms against the dataset, the results were striking:

Experimental Results Rank

Every algorithm selected by DoMO performed in the top tier of accuracy. This proves the ontology doesn't just find any algorithm; it finds the best ones without requiring the user to run a single experiment beforehand.

Why This Matters (Takeaway)

DoMO addresses the "Cold Start" problem in machine learning. Instead of burning computational resources on brute-force search (AutoML), DoMO uses deductive reasoning. This approach is particularly valuable for:

  • Interdisciplinary Researchers: Scientists in medicine or economics who need advanced ML but aren't computer scientists.
  • Scalable Engineering: Automated systems that need to switch algorithms on-the-fly as data characteristics evolve.

Future Outlook & Limitations

While DoMO integrates the most famous DM frameworks (CRISP-DM), the "Domain Level" still requires manual input from experts to define specific constraints. The next frontier will likely involve LLMs (Large Language Models) automatically populating these ontology levels, creating a truly autonomous and self-evolving data science assistant.


Summary Table: DoMO Axiom Statistics Ontology Logic Metrics

Find Similar Papers

Try Our Examples

  • Search for recent studies that integrate Semantic Web technologies with AutoML frameworks to automate the CRISP-DM lifecycle.
  • Which paper first introduced the "OntoDT" ontology, and how does DoMO extend its foundational data type properties for domain-specific tasks?
  • Investigate how multi-level ontologies like DoMO are being applied to automate feature engineering in specialized fields like Bio-medicine or Financial Fraud detection.
Contents
DoMO: Democratizing Data Science through Multi-Level Ontologies
1. TL;DR
2. Background: The Gap in Intelligent Data Assistants
3. Methodology: The DoMO Architecture
3.1. 1. The Four Layers of Insight
3.2. 2. The INPUT Ontology: The User Bridge
4. Case Study: Time Series Classification (TSC)
4.1. Experimental Evidence
5. Why This Matters (Takeaway)
6. Future Outlook & Limitations