YAM++: Breaking the Manual Threshold Barrier in Ontology Matching with Machine Learning

A Generic Approach for Combining Linguistic and Context Profile Metrics in Ontology Matching

2011-01-01
DuyHoa Ngo, Zohra Bellahsene, Remi Coletta
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a machine learning-based framework for ontology matching that integrates terminological and context profile metrics using a Decision Tree model. The approach, implemented in the YAM++ system, leverages string-based, linguistic, and structural-context features to achieve the second-highest F-measure in the OAEI 2010 Conference track without requiring manual threshold tuning.

TL;DR

Ontology matching—the process of finding correspondences between different knowledge schemas—is often a manual "trial and error" process of setting weights and thresholds. The YAM++ approach, presented by researchers at LIRMM, shifts this paradigm by using a Decision Tree model to automatically combine linguistic, string-based, and context-profile metrics. The result is a highly flexible system that achieved SOTA-level performance in the OAEI 2010 campaign.

The Problem: The Heterogeneity Headache

In a perfect world, two ontologies describing "Books" would use the same labels and structures. In reality, one might use MscThesis while another uses Ms.dissertation. Even worse, some ontologies use randomized IDs like dzqndbzq for important classes.

Existing matchers usually fall into two traps:

  1. Dependency on Terminological Similarity: They fail when names are obfuscated.
  2. Rigid Manual Tuning: They require experts to decide that "Levenstein similarity should have a weight of 0.7," which rarely works across different datasets.

Methodology: A Three-Layered Intelligence

1. Advanced Terminological Analysis

YAM++ doesn't just look at strings; it decomposes them. Using a Morphological Analysis algorithm (Algorithm 1), it uses WordNet to find the base forms of tokens.

  • Example: It recognizes that "published" (adjective) and "publishing" (verb) share the root "publish," bridging the gap where simple string matchers fail.

2. Contextual Profiling (Beyond the Name)

When names are missing or random, YAM++ looks at the "neighborhood." The authors define three types of profiles:

  • Individual Profile: Local labels and comments.
  • Semantic Profile: Information inherited from sub-classes and restricted properties.
  • External Profile: Real-world data extracted from ontology instances (A-Box).

Overall Architecture Logic Fig 1: Illustration of how Semantic and Individual profiles are constructed from neighboring nodes.

3. The Decision Tree "Brain"

Instead of a linear weighted sum, YAM++ uses a Decision Tree. This allows for non-linear logic: "If the Name similarity is high, trust it; if it's low but the Context similarity is very high, it’s still a match." This automated logic eliminates the "Threshold setting" problem found in systems like Falcon or ASMOV.

Performance and Results

The system was put to the test in the OAEI 2010 campaign, specifically on the Benchmark and Conference tracks.

  • Precision Mastery: In benchmark tests, YAM++ maintained a precision of nearly 1.0, ensuring that the discovered mappings were almost always correct.
  • Recall Improvement: As shown in the ablation-style analysis (Fig 2), adding context profiles significantly boosted recall in scenarios where naming was randomized.
  • Competitive Edge: In the Conference track, YAM++ reached the 2nd position with an F-measure of 0.61, outperforming most participants while requiring less manual intervention.

Performance Table Fig 2: Comparison of F-measure across OAEI 2010 participants.

Critical Insight: Why it Works

The brilliance of YAM++ lies in its feature selection strategy. By using Pearson’s Correlation to select the most representative metrics for each group (Name, Label, Context), the authors ensure that the Decision Tree isn't overwhelmed by redundant data. This selective approach makes the learning process both efficient and robust to the domain shift (from Benchmark training data to Conference test data).

Conclusion and Future Outlook

YAM++ demonstrates that machine learning can effectively replace human intuition in ontology matching. While it excels at linguistic and context profiling, the authors acknowledge the next frontier: integrating deeper structural and semantic reasoning (e.g., consistency checking) to further refine the quality of matches. For developers building knowledge graphs, the takeaway is clear: don't just match strings—match the context.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Deep Learning or Transformers for automated metric combination in ontology matching to compare against Decision Tree methods.
  • Which paper first introduced the "Virtual Document" or "Context Profile" concept in ontology alignment, and how has this paper expanded that concept into Semantic and External profiles?
  • Explore research that applies machine learning-based ontology matching techniques to large-scale Knowledge Graph integration or Linked Data tasks.
Contents
YAM++: Breaking the Manual Threshold Barrier in Ontology Matching with Machine Learning
1. TL;DR
2. The Problem: The Heterogeneity Headache
3. Methodology: A Three-Layered Intelligence
3.1. 1. Advanced Terminological Analysis
3.2. 2. Contextual Profiling (Beyond the Name)
3.3. 3. The Decision Tree "Brain"
4. Performance and Results
5. Critical Insight: Why it Works
6. Conclusion and Future Outlook