DMOM: Scaling Ontology Matching via Data Mining and Feature Selection

Data mining-based approach for ontology matching problem

2020-01-03
Hiba Belhadi, Karima Akli-Astouati, Youcef Djenouri, Jerry Chun-Wei Lin
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces DMOM (Data Mining for Ontology Matching), a framework designed to identify correspondences between ontology instances by treating the matching process as a feature selection problem. By utilizing Frequent Itemset Mining (FIM), the method efficiently prunes high-dimensional data properties, achieving SOTA performance on OAEI and DBpedia datasets.

TL;DR

The explosion of the Semantic Web has made ontology matching—the process of finding identical entities across different databases—computationally "expensive." DMOM (Data Mining for Ontology Matching) shifts the focus from complex alignment algorithms to a smart pre-processing step: Feature Selection. By using Frequent Itemset Mining (FIM), DMOM identifies the most relevant data properties, cutting through the noise of high-dimensional data like DBpedia to deliver a 300% to 800% speedup while hitting near-perfect accuracy.

The Scalability Wall in Ontology Matching

As we move toward a "Web of Data," ontologies are becoming massive. Consider DBpedia: it contains over 4.2 million instances and nearly 3,000 distinct properties. A naive matching approach that compares every instance and every property leads to a staggering comparisons.

Current SOTA tools like RiMOM and EIFPS struggle when property counts exceed 10%. As shown in the paper's motivation, their runtimes skyrocket as the dimensionality of the data properties increases. The bottleneck isn't just the number of instances; it's the redundant and sparse properties that add noise to the similarity calculations.

Performance Gap in Prior Work

Methodology: Feature Selection as the Silver Bullet

The core philosophy of DMOM is: Why compare 3,000 properties when only 10 define the entity? The framework introduces a two-stage pipeline:

  1. Feature Selection: Pruning the properties of each ontology independently.
  2. Matching Process: Comparing instances using only the reduced, high-value feature set.

The Winning Strategy: Frequent Itemset Mining (FIM)

While the authors tested Exhaustive and Statistical methods, the FIM Strategy proved superior. It treats each instance as a "transaction" and its properties as "items."

  • Mining: Uses the SSFIM (Single Scan) algorithm to find properties that appear frequently together.
  • Pruning: A novel "coverage" function ensures we keep the minimal set of properties that describe the maximum number of instances (following the Minimum Description Length principle).
  • Selection: Only properties with a high probability of appearing in these "frequent clusters" are used for the final alignment.

Overall DMOM Framework

Experimental Battleground: OAEI & DBpedia

The researchers didn't just test on small benchmarks; they went after the "big fish."

1. Accuracy (F-measure)

On the OAEI "Person" and "IIMB" tracks, the FIM strategy achieved an F-measure of 0.92 to 1.0. Compared to the Exhaustive strategy (which often hovered around 0.75), FIM's ability to see through "random" property distributions allowed it to find higher-quality matches.

2. Speed and Efficiency

The real triumph was on DBpedia. When handling 1,000,000 matchings:

  • Baseline (RiMOM/EIFPS): ~900+ seconds.
  • DMOM (FIM): 111 seconds.

This efficiency comes from the fact that DMOM's runtime stabilizes even as property sizes grow, whereas traditional tools scale poorly.

Runtime Comparison on DBpedia

Academic Insight: Why it Works

The success of FIM-based matching highlights an Inductive Bias in real-world ontologies: Information is not distributed uniformly. In most ontologies, a small subset of properties (like 'label', 'name', or 'birthDate') carries the vast majority of the discriminative power. By using frequent pattern mining, DMOM automatically "learns" which attributes are the essential identifiers of a domain without requiring human-in-the-loop expert rules.

Conclusion & Future Look

DMOM proves that in the era of Big Data, pre-processing is as important as the algorithm itself. By stripping away the 99% of "noise" properties in large ontologies, we can make real-time Semantic Web integration possible.

Limitations: The current framework relies on specific frequency thresholds. Future work might involve Deep Feature Embeddings to handle semantic differences where properties are frequent but named differently (e.g., has_name vs label).


Editor's Note: This paper is a must-read for anyone dealing with Knowledge Graph alignment and Record Linkage in high-dimensional spaces.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply Deep Learning or Embeddings to further enhance the "Matching Process" stage within the DMOM framework.
  • What are the foundational papers for the SSFIM (Single Scan Frequent Itemset Mining) algorithm, and how does it optimize memory compared to FP-Growth?
  • Explore how feature selection techniques from DMOM can be adapted for cross-lingual ontology matching or schema-level alignment in the Linked Open Data cloud.
Contents
DMOM: Scaling Ontology Matching via Data Mining and Feature Selection
1. TL;DR
2. The Scalability Wall in Ontology Matching
3. Methodology: Feature Selection as the Silver Bullet
3.1. The Winning Strategy: Frequent Itemset Mining (FIM)
4. Experimental Battleground: OAEI & DBpedia
4.1. 1. Accuracy (F-measure)
4.2. 2. Speed and Efficiency
5. Academic Insight: Why it Works
6. Conclusion & Future Look