Beyond Statistical Significance: Ranking Association Rules via Ontology Knowledge Mining
Ontology Knowledge Mining Based Association Rules Ranking
This paper introduces an ontology-based interestingness measure for ranking Medical Association Rules (ARs), specifically within the mammographic domain. By applying a k-medoid hierarchical clustering approach to the "Mammo" ontology, the authors calculate semantic dissimilarity to prioritize rules that link items from disparate knowledge categories.
TL;DR
In medical data mining, finding a "strong" rule is easy, but finding a "useful" one is hard. This paper presents a novel framework that uses hierarchical conceptual clustering of medical ontologies to rank Association Rules (ARs). By focusing on the dissimilarity between items (e.g., linking a symptom to a diagnosis rather than two similar symptoms), the system surfaces the most "interesting" insights for clinicians.
Problem & Motivation
Association Rule Mining (ARM) is a staple for discovering correlations in large databases. However, in the medical field, clinicians face the "Rule Explosion" problem. Traditional metrics like Support and Confidence only tell you if a relationship is frequent—not if it's insightful.
The authors identify a specific gap: current subjective measures often use simple "path-counting" in ontologies. These methods are limited because they treat all branches of an ontology tree equally. Domain experts, however, value "distal" relationships—rules that bridge different conceptual domains, such as clinical features and radiological findings.
Methodology: The Core
The paper’s innovation lies in Ontology Knowledge Mining (OKM). Instead of using the ontology as a static tree, they transform it into a hierarchy of clusters that represent distinct knowledge topics.
1. The Semantic Context Distance
The authors define a distance measure based on the "Context" of a concept. A concept's context includes its subsumption (is-a) and associative (related-to) relationships. The distance is calculated as follows:
2. Hierarchical Divisive Clustering
Using an iterative k-medoid algorithm, the ontology is broken down into clusters. Each cluster is represented by a "medoid" (a central concept). While concepts within a cluster are "similar," the distance between clusters defines the novelty of a rule.
Figure 1: Extract of the hierarchical clusters generated from the Mammo ontology, featuring medoids like 'Diagnosis' and 'Mass Shape'.
3. Measuring Rule Interestingness
A rule's interest score is the average semantic distance between its items. If items and belong to the same cluster (e.g., two different names for the same symptom), the score is low. If they span different branches (e.g., a specific age group and a malignant diagnosis), the score is high.
Experiments & Results
The study used a real-world dataset of 1,000 patients from Charles Nicolle Hospital. By extracting 1,177 rules via the Apriori algorithm, the authors applied their semantic filter.
Key Performance Data:
- Clustering Quality: The hierarchical approach achieved high Precision and Recall across 5 levels, ensuring that semantic categories like 'Anatomical_entities' and 'Diagnosis' remained distinct.
- Pruning Efficacy: By adjusting the threshold (), the system can control the "noise" presented to the doctor.
- Expert Validation: Rules ranked high by the system (Distance = 1.0) were confirmed by experts to be highly predictive, such as [Parallel_mass] → [Benign Diagnosis].
Figure 2: Scaling the number of "Interesting Rules" based on the Semantic Threshold .
| Rule Detail | Semantic Interpretation | Interest Score |
|---|---|---|
| mass_oval → benign_diagnosis | Oval mass is highly predictive of benign lesion | 1.0 |
| breast_pain → mastalgia | Trivial (same category) | Low |
Critical Analysis & Conclusion
Takeaway
The shift from "concept-to-concept" distance to "cluster-to-cluster" distance is a significant step forward. It reflects a more human-like understanding of knowledge categories, where the most valuable insights are those that connect different "silos" of medical data.
Limitations
The primary bottleneck of this approach is its dependence on a well-structured ontology. In domains where ontologies are sparse or poorly maintained, the clustering algorithm may fail to define meaningful boundaries, leading to poor ranking.
Future Work
The authors suggest that future iterations could include weights for different types of associative relations, further refining what "dissimilarity" looks like in a clinical context.
