Ontology-Based Classification: Achieving High Accuracy Without Training Data
Ontology-Based Classification – Application of Machine Learning Concepts Without Learning
The paper introduces an Ontology-Based Classification method for categorizing education offers into innovation clusters. By utilizing a domain-specific thesaurus to construct "centrality sets" rather than training on labeled data, the approach achieves a high accuracy of 87% (Naïve Bayes) and 85.5% (Cosine Similarity) in a "learning without learning" paradigm.
TL;DR
What do you do when you have 30,000 documents to classify into 9 categories, but no reliable training labels and zero budget for manual annotation? This paper presents a pragmatic breakthrough: Ontology-Based Classification. By leveraging a pre-existing domain thesaurus, the author constructs a classification model directly from structured knowledge, achieving an impressive 87% accuracy without a single step of traditional machine learning "training."
Problem & Motivation: The "Small Data" Wall
In the academic world, we often assume the availability of neatly labeled datasets like ImageNet or MNLI. In the real world—specifically the "Weiterbildungsdatenbank Berlin-Brandenburg" (WDB-BB)—the situation is grimmer:
- Imperfect Data: Existing classifications were done by institutions with varying standards, leading to inconsistent labels.
- Resource Constraints: Manually re-classifying thousands of education offers was financially and logistically impossible.
- The Goal: Map education offers to nine specific "Innovation Clusters" (e.g., Health, Energy, ICT) to support regional economic strategy.
The researcher’s insight was profound: if we already have a semantic thesaurus that defines the relationships between professional terms, why not use that structure as our model's "weights" instead of trying to learn them from scratch?
Methodology: The "Virtual Document" Trick
The core of the method lies in converting static ontological data into a dynamic classification tool through two primary steps:
1. Extracting Concept Centrality
The author treats each innovation cluster as a node in the thesaurus. By traversing the graph and collecting related terms (sub-concepts, synonyms, and relations), the system counts how often a concept is encountered.
- Centrality: A term that appears frequently during the traversal is "more central" to the cluster.
- Virtual Document: These centrality counts are treated as term frequencies in a "virtual document" that perfectly represents the cluster.

2. Matching via Standard Metrics
Once the "virtual documents" for the nine clusters were created, the classification of a new education offer became a simple similarity problem. The author compared:
- Cosine Similarity: Measuring the angle between the offer's vector and the cluster's centrality vector.
- Naïve Bayes: Calculating the probability of a cluster given the terms in the offer.
Experiments & Results: Performance without Training
The results challenge the necessity of deep training in specialized domains. By simply cleaning the centrality sets (removing non-indicative general terms) and ignoring low-centrality noise, the system achieved professional-grade performance.
| Cluster | Precision (Naïve Bayes) | Recall (Naïve Bayes) |
|---|---|---|
| Health | 89.3% | 97.9% |
| ICT | 87.6% | 85.6% |
| Metal | 98.2% | 81.5% |
| Total Accuracy | 87% | - |

One critical finding was the skewed distribution of data. Clusters like "Plastics/Chemistry" had so few samples that results were statistically insignificant. However, for the major clusters, the system provided high-confidence classifications that were immediately ready for production use.
Critical Analysis & Conclusion
The Takeaway
The "Ontology-Based Classification" approach is a masterclass in Knowledge Engineering. It proves that in domain-specific tasks where vocabulary is well-defined (like job markets or medical fields), the structure of the language itself is a more efficient carrier of information than a thousands-of-samples training set.
Limitations & Future Work
- Cold Start for Ontologies: This method requires a high-quality thesaurus. If the domain knowledge isn't already modeled, the "learning" bottleneck simply shifts from label-collecting to ontology-building.
- Syntactic Limitations: The current model is purely numerical/syntactical and cannot disambiguate terms that have different meanings in different contexts unless explicitly modeled in the thesaurus.
In the future, integrating Large Language Models (LLMs) to handle the disambiguation while using the Thesaurus Centrality as a grounding mechanism could lead to even more robust "training-free" systems.
Editor's Note: This paper serves as a reminder that the best AI solution isn't always the biggest neural network; sometimes, it's the smartest use of existing knowledge.
