Beyond Code: Automating Business Logic with Ontology-Driven Machine Learning

Automating Implementation of Business Logic of Multi Subject-Domain IS on the Base of Machine Learning, Data Programming and Ontology-Based Generation of Labeling Functions

2021-01-01
Maksim G. Shishaev, Pavel Lomov
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a model-centric architecture for Multi Subject-Domain Information Systems (MSIS) that implements business logic using Machine Learning (ML). It leverages a data programming approach with an innovative technology for the automated generation of labeling functions (LFs) based on domain ontologies to accelerate the creation of training sets.

TL;DR

Implementing business logic in complex, large-scale systems is often a losing battle against changing requirements. This paper proposes a paradigm shift: instead of hard-coding rules, we should use Machine Learning as a universal approximator for business logic, fueled by a Data Programming approach that automatically generates training data from existing Domain Ontologies.

Background: The Crisis of Multi-Domain Systems

In regional management, Information Systems (MSIS) face a "heterogeneity trap." Users, data sources, and functional requirements vary wildly across strategic and operational levels. Traditional methods like Domain-Driven Design (DDD) or Model-Driven Architecture (MDA) simply shift the burden of formalization from programmers to model designers—they don't eliminate the manual labor.

The authors argue that we should treat business logic as a "black box" that approximates responses to input data. To make this work, we need Machine Learning, but we lack the high-quality labeled data to train it for specific niche domains.

The Core Insight: Data Programming & Ontologies

The paper introduces a Model-Centric Architecture based on Weak Supervision. Instead of experts labeling thousands of data points, they write Labeling Functions (LFs)—simple code snippets that encapsulate heuristics.

The breakthrough here is the Automated Generation of LFs. By extracting knowledge from OWL (Web Ontology Language) ontologies, the system can automatically create LFs that recognize entities, relationships, and context without manual coding.

Model-Centric Architecture based on Data Programming

Methodology: How the Generation Works

The system utilizes a meta-model to transform ontology fragments into Labelling Patterns (LP).

  1. Universal Generators: Use basic OWL structures (Classes, Individuals, Object Properties) to create general lookup rules.
  2. Domain-Specific Generators: Leverage Content Ontology Design Patterns (CODP). For instance, an "Object-Role" pattern can generate syntactic patterns like "X is used as Y" or "The function of X is Y."
  3. The Pipeline: These patterns are wrapped into Python-based LFs compatible with the Snorkel framework, which then reconciles conflicting labels using a generative model to create a "noise-aware" training set.

Labeling Function Generation Scheme

Experimental Validation: The Arctic Activity Case

The authors tested this on a Named Entity Recognition (NER) task involving "Economic activity in the Arctic"—a domain where standard training sets (like Wikipedia) fail miserably.

  • The Problem: A standard multilingual spaCy model performed poorly (Precision: 0.12) because it didn't understand Arctic-specific entities.
  • The Solution: LFs generated from a domain ontology automatically labeled a specialized corpus.
  • The Result: The precision jumped to 0.97. While recall (0.42) shows room for improvement (likely requiring more text data), the leap in accuracy proves that ontologies can effectively "teach" ML models niche logic.

Critical Analysis & Future Outlook

This approach elegantly bridges the "statistical" and "cognitive" gap in AI. However, it isn't a silver bullet:

  • The Data Requirement: You still need a large volume of unlabeled data.
  • Pattern Recognition Limits: The logic must be representable as pattern recognition.
  • Ontology Complexity: Creating high-quality ontologies is itself a complex task.

Conclusion

The transition to "Programming without Programming" is closer than it appears. By converting formalized domain knowledge into the "fuel" for machine learning, we can build information systems that adapt as fast as the domains they represent. For industry practitioners, the message is clear: invest in your knowledge graph (ontology), as it will eventually write your software's business logic.

Find Similar Papers

Try Our Examples

  • Examine recent state-of-the-art frameworks that integrate Knowledge Graphs with the Snorkel weak supervision pipeline for automated data labeling.
  • Which seminal papers first established the transition from Model-Driven Architecture (MDA) to Model-Centric AI, and how does this paper's ontology-to-LF mapping compare?
  • Explore the application of ontology-driven labeling functions in multi-modal fields such as satellite imagery or industrial IoT for regional management.
Contents
Beyond Code: Automating Business Logic with Ontology-Driven Machine Learning
1. TL;DR
2. Background: The Crisis of Multi-Domain Systems
3. The Core Insight: Data Programming & Ontologies
4. Methodology: How the Generation Works
5. Experimental Validation: The Arctic Activity Case
6. Critical Analysis & Future Outlook
7. Conclusion