Ontology-Based EMD: Bridging Semantic Gaps in Web Document Classification
Ontology-Based Automatic Classification and Ranking for Web Documents
This paper introduces an ontology-based automatic classification and ranking framework for web documents. It leverages Earth Mover's Distance (EMD) and WordNet-based similarity to map documents to categories, achieving a total recall of 93.9% and precision of 82.1% across 15 categories.
Executive Summary
TL;DR: This paper presents a robust framework for classifying and ranking web documents by replacing traditional training-heavy classifiers with semantic ontologies. By utilizing Earth Mover's Distance (EMD) and WordNet, the method achieves high precision without the "laborious" process of manual labeling, while simultaneously providing a ranking mechanism for categorized results.
Context: Positioned as an evolution of early 2000s semantic web research, this work addresses the shift from purely statistical methods (like early KNN or Naive Bayes) to knowledge-driven architectures. It is a refinement of ontology-based classification that focuses on automation and practical sorting.
Problem & Motivation: The Limits of "Word-as-Token"
Traditional machine learning algorithms for text, such as SVM and KNN, treat words as discrete tokens. This leads to two critical failures:
- Semantic Blindness: They cannot recognize that "soccer" and "football" are related unless both appear frequently in training data.
- Training Fatigue: If the categories change, thousands of documents must be re-labeled to re-train the model.
The authors' insight is that if we can model a Category as a structured tree of concepts (an Ontology), we can classify documents in real-time by measuring how much "work" it takes to transform a document's vocabulary into that ontology's structure.
Methodology: EMD and Semantic Transport
The core of the methodology lies in the dual approach of Ontology Pretreatment and Semantic Measuring.
1. Automated Ontology Augmentation
Instead of building ontologies from scratch, the system takes an existing hierarchy (like DMOZ) and:
- Disambiguates: Uses context to determine the correct sense of a concept.
- Augments: Pulls hypernyms and synonyms from WordNet to broaden the "semantic net" of each category.
- Weights: Assigns weights based on hierarchy depth (), assuming higher-level nodes are more abstract and defining.
2. The Earth Mover's Distance (EMD)
The paper treats classification as a "transportation problem." A document is a distribution of weighted terms; an ontology is a distribution of concepts. EMD measures the minimum cost to "move" the document's weights to match the ontology's concepts based on their semantic distance in WordNet.
(Note: Refer to the paper's description of the weight graph G={d,o,S} for the optimization flow.)
Experiments & Results
The authors tested their method against a large dataset of 10,842 web documents across 15 categories.
Performance vs. Traditional SOTA
The results prove that semantic-aware classification holds a distinct advantage over distance-based KNN:
| Method | Total Recall | Total Precision |
|---|---|---|
| Our Algorithm (EMD) | 93.9% | 82.1% |
| Adaptive KNN | 86.9% | 79.3% |
| SVM | 93.5% | 84.2% |

The EMD method achieved a recall of 96.2% in the "Sports" category, demonstrating that for domains with rich, well-defined terminology, ontologies are highly effective.
Critical Analysis & Conclusion
Takeaways
The marriage of WordNet (as a linguistic backbone) and EMD (as a mathematical distance measure) creates a classifier that understands intent rather than just frequency. The ranking mechanism, while simple, provides an essential layer for user experience in document navigation.
Limitations
- Computational Cost: EMD (the transportation problem) is significantly more computationally expensive than a simple KNN pass.
- Linguistic Dependency: The performance is heavily tied to WordNet’s coverage. If a domain uses slang or new technical jargon not in WordNet, the similarity score collapses.
Future Outlook
The authors suggest moving toward Weighted Word Senses—using specific meanings rather than raw words. In the modern era, this logic is the precursor to Dense Vector Embeddings, where we now use cosine similarity in latent space instead of manual EMD on ontology graphs.
