Bridging the Cold-Start Gap: Semantic Similarity in Sparse Data Mining
Using Ontology-Based Similarity Measures to Find Training Data for Problems with Sparse Data
This paper introduces three ontology-based similarity measures—Graph Distance, Direct Neighbors (feature-based), and Tree Distance—designed to facilitate data mining in sparse data scenarios. By mapping unstructured metadata to semantic concepts, the method aggregates data from similar sources to train predictive models for new entities.
TL;DR
When a new customer joins an online shop, there is no history to train a recommendation model. This paper proposes using ontologies—formal structures of knowledge—to compute similarities between new and existing entities based on their metadata. By comparing concepts like "Adidas" and "Nike" or "Sandal" and "Beach Shoe" through their positions in a knowledge graph, the system can aggregate training data from similar "neighbors" to provide immediate, accurate predictions.
The Sparsity Struggle: Why Keywords Aren't Enough
In the world of online retail, data is king. But what happens when you have a "cold start"? A new user registers, provides a brief bio or interest list, and has zero purchase history.
Standard data mining falls short here because:
- Syntactic Mismatch: A user interested in "Sprints" and a user interested in "Running" might be seen as completely different by simple keyword matchers.
- Context Blindness: Traditional algorithms don't know that a "Suit" and a "Tie" belong together unless they see them bought together millions of times.
The authors argue that we should use the Semantic Web approach to inject "common sense" into these algorithms.
Methodology: Mapping Text to Meaning
The authors propose a three-step pipeline to transform messy text into a similarity score:
- Term Extraction: Cleaning metadata by removing stop words and stemming.
- Concept Mapping: Linking extracted terms to nodes in an ontology.
- Semantic Measurement: Using the graph structure to calculate how "close" two entities are.

Three Flavors of Similarity
- Graph Distance (GD): Measures the shortest path between nodes. It’s fast but sometimes "blind" to the type of relationship.
- Direct Neighbors (DN): A feature-based approach. It looks at the "friends" of a concept. If two terms share many neighbors in the graph, they are likely similar.
- Tree Distance (TD): Specifically looks for a "Common Ancestor." While theoretically sound for hierarchies, it proved computationally expensive in practice.
Experimental Insight: The Fashion Ontology
To prove their point, the authors built a massive Fashion Ontology with over 1,700 brands and 700 categories.

The results were telling. The Direct Neighbors approach was the "Goldilocks" solution. It correctly identified that:
{"adidas", "t-shirt"}is highly similar to{"nike", "t-shirt"}.{"sandal"}is semantically related to{"shoe", "summer", "beach"}even though they share zero words.
Conversely, Graph Distance struggled because, in a highly connected graph, everything is "close" to everything else, leading to diluted similarity scores.
Critical Analysis & Conclusion
This work highlights a vital shift from pure "black box" machine learning toward Semantic Machine Learning.
The Takeaway: If you are dealing with sparse data, don't just wait for more data—use existing expert knowledge (Ontologies). The feature-based "Direct Neighbors" measure is a robust way to do this without the exponential time costs of searching for common ancestors.
Limitations: The reliance on a manually curated or auto-crawled ontology means the system is only as good as its knowledge base. Furthermore, the 3-hour runtime for Tree Distance indicates that hierarchical searches need significant optimization (perhaps via indexing) before they can be used in real-time retail environments.
Looking ahead, the integration of these semantic measures into the loss functions of Deep Learning models could represent the next frontier in solving the cold-start problem.
