SemSim: Reaching SOTA in Semantic Similarity via Hybrid LSA and Linguistic Fusion
Robust semantic text similarity using LSA, machine learning, and linguistic resources
The paper introduces SemSim, a robust system for Semantic Textual Similarity (STS) that achieved top rankings in *SEM 2013 and SemEval-2014. It utilizes a hybrid word similarity model combining Latent Semantic Analysis (LSA) with WordNet knowledge, integrated through an unsupervised term alignment algorithm and supervised regression models.
Executive Summary
TL;DR: The SemSim system represents a high-water mark in the "pre-Transformer" era of NLP, achieving 1st place in major *SEM and SemEval competitions. It succeeds by fusing the statistical "intuition" of Latent Semantic Analysis (LSA) with the formal structure of WordNet and the real-time context of Web-based dictionaries.
Field Positioning: This work bridges the gap between pure distributional semantics and knowledge-based reasoning. It proves that even before the dominance of BERT, carefully engineered alignment and external resource integration could achieve human-level correlation in judging textual equivalence.
The Core Problem: Why Words Are Not Enough
Measuring Semantic Textual Similarity (STS) is deceptively simple: are "A man is dancing" and "A person is performing rhythmic movements" the same? To a computer, the word vectors for "dancing" and "performing" might be distant.
The authors identify two fatal flaws in prior SOTA:
- Contextual Poverty: Short texts lack the statistical density for Bag-of-Words to work.
- The OOV Wall: Names, slang, and new technical terms (e.g., "Google" as a verb) break static vocabularies.
Methodology: The "Knowledge-Augmented Distributional" Approach
The genius of SemSim lies in its three-layered architecture.
1. Robust Word Similarity (The Engine)
The system doesn't just use LSA; it uses a 3-billion-word high-quality corpus from the Stanford WebBase. By applying SVD with varying window sizes ( for concept similarity and for relation similarity), they captured different semantic dimensions.
Crucially, they hybridize LSA with WordNet using a boosting formula: Where is the path distance in WordNet. This ensures that even if two words rarely co-occur, their structural relationship (synonyms/hypernyms) pulls them together.
2. Term Alignment (The Logic)
Rather than a flat vector comparison, SemSim aligns terms between sentences. It uses Information Content (IC) to weight words, ensuring that "cardiologist" carries more weight than "doctor."
Figure 1: High-level architecture of the SemSim system showing the flow from English/Spanish input to the core similarity model.
3. Handling the "Unknowable": OOV and Slang
When the system encounters a word like "skimp" or "braless" (OOV), it doesn't give up. It crawls Wordnik and Urban Dictionary in real-time, using the definitions of these words as a proxy for the words themselves.
Experiments & Results
SemSim was the undisputed champion of the SemEval-2014 Cross-Level tasks.
- LSA Performance: On TOEFL synonyms, SemSim reached 96.2% accuracy, beating Google's Word2Vec (84.8%).
- Cross-Level Mastery: The system excelled at comparing different scales—matching a single word (e.g., "confess") to an idiom (e.g., "spill the beans").
Table: SemSim rankings across various SemEval-2014 subtasks, highlighting the 1st place finishes.
Critical Insight & Conclusion
The success of SemSim provides a vital lesson for modern AI: Distributional models (like LSA or Embeddings) are powerful, but they are "deaf" to the structured logic of human language. By injecting WordNet hierarchies and web-search results, the authors created a system that doesn't just calculate co-occurrence—it understands relationship and context.
Limitations: The system relies heavily on the quality of external APIs (Google Translate, Bing Research). If the Urban Dictionary definition is a "joke" (e.g., the Programmer/Caffeine example in the paper), the similarity score tanks.
Future Outlook: SemSim sets the stage for "Retrieval-Augmented Generation" (RAG) concepts by showing that external knowledge fetching is the only way to handle the evolving long-tail of human language.
