Bridging the Gap: Scalable Semantic Similarity via Web-Contextualized Ontologies

Ontology-driven web-based semantic similarity

2009-10-13
David Sánchez, Montserrat Batet, Aïda Valls, Karina Gibert
Summary
Problem
Method
Results
Takeaways
Abstract

The paper proposes a novel method for computing semantic similarity between concepts by combining ontological knowledge with Web-scale statistics. It introduces the "Web-based Contextualized Information Content" (CICT_IR), which leverages search engine hit counts constrained by taxonomic subsumers to outperform classical measures like Resnik, Lin, and Jiang & Conrath.

Executive Summary

TL;DR: This paper introduces a scalable framework to measure semantic similarity by leveraging the vast statistics of the Web while filtering out "noise" using taxonomic structures. By forcing word co-occurrences with their ontological ancestors (e.g., "dog" in the context of "mammal"), the authors solve the twin problems of data sparseness in small corpora and ambiguity in massive ones.

Positioning: This work is a significant "architectural bridge." It takes classical Information Content (IC) theories—pioneered by Resnik and Lin—and adapts them for the "Wild West" of the Web, transforming unsupervised statistical tools into semantically aware similarity engines.

The "Generalization" Paradox: Why Big Data Isn't Enough

Calculating how similar "apple" is to "pear" seems simple, but for AI, it’s a minefield:

  1. Structural Issues: Path-based measures rely on every link in an ontology being "equal," which isn't true (the distance between "Entity" and "Object" is semantically larger than "Dog" and "Labrador").
  2. Corpus Issues: Using small, manually tagged corpora (like SemCor) is highly accurate but covers less than 13% of WordNet.
  3. The Web's Chaos: Using the Web solves the coverage problem, but raw hit counts are misleading. For example, "dog" (the animal) appears millions of times, but so does "hot dog" (food) or "dog" (a mechanical part). Without context, the statistics are poisoned.

Methodology: Contextualization as a Semantic Filter

The core innovation is Contextualized Information Content (CICT_IR). Instead of querying the Web for a term a in isolation, the authors query for a AND LCS(a, b), where LCS is the Least Common Subsumer (the closest shared ancestor in a hierarchy).

The Intuition

If you want to know the probability of "terrier" in the sense of an animal, you query terrier AND dog. This achieves two things:

  • Taxonomic Coherence: It ensures the descendant term can never be "more probable" than its ancestor, maintaining the mathematical requirements of Information Theory.
  • Disambiguation: It automatically filters out documents where "terrier" might refer to something non-canine.

Concept Probability Problem Figure 1: Comparison of word hit counts in an ontology showing how uncontextualized raw counts (like 'dog' vs 'mammal') can lead to non-monotonic IC values.

The authors specifically applied this to Symmetric Conditional Probability (SCP) and Pointwise Mutual Information (PMI), transforming these traditionally "blind" statistical measures into "ontology-guided" experts.

Experiments: Superior Statistics

The authors tested their methods against the gold standard: human similarity judgments.

Key Competitive Results:

  • Raw Web Statistics: Measures using raw hit counts performed poorly (Correlation ~0.35 - 0.48).
  • The Proposed Fix: By adding the taxonomic context (CICT_IR), performance surged. The SCP_CICT_IR version reached a correlation of 0.739.
  • Benchmark Comparison: This result is remarkably close to the 0.79 correlation achieved by Resnik's original model, which required an expensive, hand-labeled corpus.

Experimental Results Table 2: Quantitative comparison showing the significant jump in correlation with human ratings when moving from IR (raw web) to CIC_T_IR (contextualized web) measures.

Critical Insight & Conclusion

Takeaway

The genius of this approach is its scalability. By using search engine "hit counts" as a proxy for social-scale word usage and using the ontology as a logical "mask," we can compute high-quality semantic similarity for any domain (medicine, chemistry, law) as long as a basic taxonomy exists.

Limitations

  • Search Engine Reliability: As noted, Google’s hit counts can be inconsistent; the method relies on a stable indexing tool (like Bing).
  • Taxonomy Dependency: While it reduces corpus dependency, it still requires a high-quality taxonomy.

Future Outlook

This method paves the way for "live" semantic similarity—systems that evolve as the Web's content changes, without ever needing a human to manually tag a new corpus. It is particularly promising for unsupervised clustering and privacy-preserving data masking.

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend Web-based semantic similarity measures to include non-taxonomic relationships like meronymy or holonomy.
  • Which paper first introduced the "Google Similarity Distance" (GSD), and how does the CICT_IR method's use of the LCS contrast with GSD's normalized distance approach?
  • Find studies applying ontology-driven semantic similarity measures to privacy-preserving data anonymization or microaggregation in the biomedical domain.
Contents
Bridging the Gap: Scalable Semantic Similarity via Web-Contextualized Ontologies
1. Executive Summary
2. The "Generalization" Paradox: Why Big Data Isn't Enough
3. Methodology: Contextualization as a Semantic Filter
3.1. The Intuition
4. Experiments: Superior Statistics
4.1. Key Competitive Results:
5. Critical Insight & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook