DoMO: Democratizing Data Science through Multi-Level Ontologies
Domain-Oriented Multilevel Ontology for Adaptive Data Processing
The paper introduces DoMO (Domain-oriented Multi-level Ontology), a comprehensive semantic framework designed to bridge the gap between complex data mining algorithms and non-expert researchers. By integrating existing ontologies like DMOP and OntoDT into a four-level architecture, it enables automated, adaptive selection of data processing workflows and algorithms based on specific dataset characteristics.
TL;DR
The complexity of modern data mining (DM) often leaves non-expert researchers lost in a sea of algorithms. DoMO (Domain-oriented Multi-level Ontology) is a breakthrough framework that uses semantic meta-mining to automatically recommend the best data processing workflows. By mapping dataset metadata to algorithm strengths, it transforms data science from a "trial-and-error" black box into a logical, queryable knowledge base.
Background: The Gap in Intelligent Data Assistants
The "No Free Lunch" theorem dictates that no single algorithm is best for every task. In the world of Data Mining (DM), selecting the right workflow (Preprocessing -> Algorithm -> Evaluation) requires deep expertise. While Meta-learning aims to solve this by learning from past experiments, it often treats algorithms as black boxes.
Semantic Meta-mining offers a superior alternative: it uses ontologies to understand the internal mechanisms of algorithms and the intrinsic properties of data. However, previous ontologies like DMOP or OntoDM were either too narrow in scope or too complex for practical use.
Methodology: The DoMO Architecture
DoMO's power lies in its multi-level structure, which organizes knowledge from abstract data types to specific implementations.

1. The Four Layers of Insight
- Upper Level: Uses OntoDT to define the fundamental DNA of data (basic types, value spaces).
- Domain Level: Where experts define what "Short," "Long," or "Noisy" means for their specific field (e.g., Time Series vs. Image Recognition).
- Application Level: A specific core ontology tailored for a domain.
- Implementation Level: The "User Interface" where natural queries are translated into SPARQL-like logic to find executable code.
2. The INPUT Ontology: The User Bridge
The researchers created a specialized INPUT Ontology. This acts as a semantic interface where a user can describe their dataset in plain terms (e.g., "Dataset with 1000 samples, high noise, 3 classes") and receive a full-stack algorithm recommendation.
Case Study: Time Series Classification (TSC)
To prove DoMO's effectiveness, the authors applied it to the CinCECGtorso dataset (ECG data).
The Workflow:
- Describe: The user identifies the dataset as
SmallTrainTSDataset,LongTSDataset, andECGTSDataset. - Query: The system searches for algorithms
suitableForthese specific constraints. - Recommend: It identified BOSS, COTE, EE, MSM_1NN, and ST as optimal candidates.
Experimental Evidence
When the authors benchmarked all available TSC algorithms against the dataset, the results were striking:

Every algorithm selected by DoMO performed in the top tier of accuracy. This proves the ontology doesn't just find any algorithm; it finds the best ones without requiring the user to run a single experiment beforehand.
Why This Matters (Takeaway)
DoMO addresses the "Cold Start" problem in machine learning. Instead of burning computational resources on brute-force search (AutoML), DoMO uses deductive reasoning. This approach is particularly valuable for:
- Interdisciplinary Researchers: Scientists in medicine or economics who need advanced ML but aren't computer scientists.
- Scalable Engineering: Automated systems that need to switch algorithms on-the-fly as data characteristics evolve.
Future Outlook & Limitations
While DoMO integrates the most famous DM frameworks (CRISP-DM), the "Domain Level" still requires manual input from experts to define specific constraints. The next frontier will likely involve LLMs (Large Language Models) automatically populating these ontology levels, creating a truly autonomous and self-evolving data science assistant.
Summary Table: DoMO Axiom Statistics

