Beyond Labeling: Solving Ontology Mapping with PU Learning and Entropy
Semi-supervised Learning Approach for Ontology Mapping Problem
This paper proposes a novel semi-supervised learning framework for ontology mapping, formulated as a Positive and Unlabeled (PU) learning problem. The method utilizes a Naive Bayesian classifier enhanced by an entropy-based artificial negative example generation technique (LGN) to identify correspondences between heterogeneous ontologies.
TL;DR
Ontology mapping is the cornerstone of the Semantic Web, yet it suffers from a "negative data scarcity" problem. This paper presents a semi-supervised approach that treats mapping as a PU Learning (Positive and Unlabeled) task. By using entropy to generate artificial negative examples, the researchers build a Naive Bayesian classifier that reliably identifies entity correspondences without requiring manually labeled non-matches.
The Core Motivation: The "Negative Label" Bottleneck
In the world of Semantic Web interoperability, finding "what matches" is intuitive for human experts, but "what doesn't match" creates an infinite search space. Existing SOTA methods usually fall into two traps:
- Linear Over-simplification: Aggregating similarity scores (lexical, structural, etc.) using manually assigned weights, which lacks adaptability.
- Supervised Dependency: Relying on large training sets containing both positive and negative samples, which are rarely available in real-world large-scale ontologies.
The authors' insight is profound: If we only have positive matches, we can treat the rest of the data as a "sea of unlabeled potential," using information theory to fish out the true negatives.
Methodology: The PU Learning Framework
The proposed workflow converts the mapping process into a high-dimensional classification problem.
1. Constructing the Feature Space
Instead of relying on a single similarity metric, the system generates a Similarity Vector for every pair of entities . This vector includes:
- Lexical similarity (WordNet-based).
- Structural/Hierarchy similarity.
- Instance-level properties.
2. The Entropy-Based Negative Generator
The "secret sauce" of this paper is the generation of artificial negative data. The authors use an entropy measure to identify which features are most "certain" or "uncertain."

- Low Entropy: Indicates a feature value that is significantly more prevalent in the unlabeled set than in the positive set—marking it as a strong candidate for a negative attribute.
- High Entropy: Indicates the distribution is similar across sets, making it a "hidden positive" candidate.
These weights are then used to synthesize an Artificial Negative Instance (), which acts as a proxy for the missing negative class.
3. Naive Bayesian Classification
With the Positive set () and the Artificial Negative (), a Naive Bayesian classifier is trained using the Bayes formula:

The classifier identifies "outliers" in the unlabeled set—those instances that look nothing like the positive matches—and labels them as negatives, effectively filtering the set until only the true matches remain.
Experiments and Benchmarking
The researchers utilized the OAEI (Ontology Alignment Evaluation Initiative) 2010 Benchmark, specifically tasks #301-#304 involving real-world biography ontologies.
Key Metrics:
- Recall & Precision: The paper emphasizes the harmonic mean (F1-score) to ensure that the artificial negatives don't lead to over-filtering (Precision) or missing true matches (Recall).
- Comparison: The approach is positioned against the ASCO algorithm and traditional SVM-based mappings, where it shows superior flexibility in the absence of a balanced training set.
Critical Insight: Why Entropy Matters
Most ontology tools treat all similarity features equally. By introducing entropy, this method recognizes that some features (like exact label matches) are highly discriminative, while others (like shared parent classes) might be noise. This Information-Theoretic Inductive Bias allows the model to learn the "structure of similarity" rather than just the "value of similarity."
Conclusion & Future Outlook
This paper successfully demonstrates that label scarcity is not a deal-breaker for ontology mapping. The use of PU learning shifts the paradigm from "finding a needle in a haystack" to "systematically removing the hay."
Limitations: The current model requires feature discretization for the Naive Bayesian approach, which might lose some nuance in similarity scores. Next Steps: Integration of multi-class classification and refining the classifier to handle continuous Gaussian distributions without discretization could further boost SOTA performance in the Semantic Web domain.
