DMOM: Scaling Ontology Matching via Data Mining and Feature Selection
Data mining-based approach for ontology matching problem
This paper introduces DMOM (Data Mining for Ontology Matching), a framework designed to identify correspondences between ontology instances by treating the matching process as a feature selection problem. By utilizing Frequent Itemset Mining (FIM), the method efficiently prunes high-dimensional data properties, achieving SOTA performance on OAEI and DBpedia datasets.
TL;DR
The explosion of the Semantic Web has made ontology matching—the process of finding identical entities across different databases—computationally "expensive." DMOM (Data Mining for Ontology Matching) shifts the focus from complex alignment algorithms to a smart pre-processing step: Feature Selection. By using Frequent Itemset Mining (FIM), DMOM identifies the most relevant data properties, cutting through the noise of high-dimensional data like DBpedia to deliver a 300% to 800% speedup while hitting near-perfect accuracy.
The Scalability Wall in Ontology Matching
As we move toward a "Web of Data," ontologies are becoming massive. Consider DBpedia: it contains over 4.2 million instances and nearly 3,000 distinct properties. A naive matching approach that compares every instance and every property leads to a staggering comparisons.
Current SOTA tools like RiMOM and EIFPS struggle when property counts exceed 10%. As shown in the paper's motivation, their runtimes skyrocket as the dimensionality of the data properties increases. The bottleneck isn't just the number of instances; it's the redundant and sparse properties that add noise to the similarity calculations.

Methodology: Feature Selection as the Silver Bullet
The core philosophy of DMOM is: Why compare 3,000 properties when only 10 define the entity? The framework introduces a two-stage pipeline:
- Feature Selection: Pruning the properties of each ontology independently.
- Matching Process: Comparing instances using only the reduced, high-value feature set.
The Winning Strategy: Frequent Itemset Mining (FIM)
While the authors tested Exhaustive and Statistical methods, the FIM Strategy proved superior. It treats each instance as a "transaction" and its properties as "items."
- Mining: Uses the SSFIM (Single Scan) algorithm to find properties that appear frequently together.
- Pruning: A novel "coverage" function ensures we keep the minimal set of properties that describe the maximum number of instances (following the Minimum Description Length principle).
- Selection: Only properties with a high probability of appearing in these "frequent clusters" are used for the final alignment.

Experimental Battleground: OAEI & DBpedia
The researchers didn't just test on small benchmarks; they went after the "big fish."
1. Accuracy (F-measure)
On the OAEI "Person" and "IIMB" tracks, the FIM strategy achieved an F-measure of 0.92 to 1.0. Compared to the Exhaustive strategy (which often hovered around 0.75), FIM's ability to see through "random" property distributions allowed it to find higher-quality matches.
2. Speed and Efficiency
The real triumph was on DBpedia. When handling 1,000,000 matchings:
- Baseline (RiMOM/EIFPS): ~900+ seconds.
- DMOM (FIM): 111 seconds.
This efficiency comes from the fact that DMOM's runtime stabilizes even as property sizes grow, whereas traditional tools scale poorly.

Academic Insight: Why it Works
The success of FIM-based matching highlights an Inductive Bias in real-world ontologies: Information is not distributed uniformly. In most ontologies, a small subset of properties (like 'label', 'name', or 'birthDate') carries the vast majority of the discriminative power. By using frequent pattern mining, DMOM automatically "learns" which attributes are the essential identifiers of a domain without requiring human-in-the-loop expert rules.
Conclusion & Future Look
DMOM proves that in the era of Big Data, pre-processing is as important as the algorithm itself. By stripping away the 99% of "noise" properties in large ontologies, we can make real-time Semantic Web integration possible.
Limitations: The current framework relies on specific frequency thresholds. Future work might involve Deep Feature Embeddings to handle semantic differences where properties are frequent but named differently (e.g., has_name vs label).
Editor's Note: This paper is a must-read for anyone dealing with Knowledge Graph alignment and Record Linkage in high-dimensional spaces.
