Beyond Keywords: Leveraging Lexical Hierarchies for Semantic Node Similarity

Measures of Semantic Similarity of Nodes in a Social Network

2014-01-01
Ahmad Rawashdeh, Mohammad Rawashdeh, Irene Díaz, Anca L. Ralescu
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a unified semantic similarity measure for social network nodes by combining Natural Language Processing (NLP) with WordNet hierarchies. It proposes the WordNet-Cosine Similarity, which transforms unstructured profile text into numerical vectors based on semantic hypernym distances to achieve more nuanced node comparison than traditional keyword matching.

TL;DR

Quantifying the similarity between individuals in a social network is a cornerstone of recommendation engines and community detection. This paper shifts the paradigm from simple keyword matching to a Unified Semantic Similarity Measure. By integrating WordNet's hierarchical structure with Cosine Similarity, the authors transform raw interests (like movie titles) into distance-based vectors, providing a much higher resolution of "relatedness" than traditional methods.

The "Identity Trap" in Social Graphs

Most social network analysis tools treat user profiles as "bags of words." If User A likes "Comedy" and User B likes "Humorous Drama," a keyword-based system might see zero overlap. This is the identity trap: the inability to recognize that distinct symbols often point to the same concept.

Prior works like the Occurrence Frequency (OF) measure or Jaccard Index focus heavily on string intersection. As shown in the authors' research, these methods often result in skewed data where users are either "identical" or "completely different," leaving no room for the nuanced "shades of grey" that define human interests.

Methodology: Mapping Language to Space

The authors propose a multi-stage pipeline to turn unstructured text into a comparable metric space:

  1. Semantic Tagging: Using the Penn Treebank tag-set to filter out irrelevant words (like conjunctions) and focus on Nouns (NN, NNP).
  2. Lexical Mapping: Every noun is fed into WordNet to retrieve its first "Synset" (set of synonyms).
  3. Hierarchy Encoding: Instead of using the word itself, they calculate the distance from the word to the root concept ([entity]). For example, "Comedy" might be 7 steps away from "entity."
  4. Vectorization: A profile becomes a vector of these hierarchical depths.
  5. Cosine Calculation: The final similarity is the cosine of the angle between two profile vectors.

Methodology Flowchart Figure 1: High-level overview of the unified similarity measure workflow.

Experimental Validation

Using a real-world dataset of 2,013 Facebook profiles (specifically focused on "Movies" interests), the researchers compared their WordNet-Cosine approach against traditional benchmarks.

Key Findings:

  • OF Measure Weakness: The Occurrence Frequency measure tended to cluster results at a similarity of 1.0, failing to provide discriminative power.
  • WordNet Distribution: The proposed method provided a much more realistic distribution of similarities, with a peak around 0.2, allowing for better ranking of potential "friends" or "interests."

Comparison of OF vs WordNet Similarity Figure 2: Contrast between Occurrence Frequency (skewed) and WordNet Similarity (distributed).

Critical Insight

The brilliance of this work lies in its simplicity: it uses the depth of a concept in a human-curated taxonomy as a proxy for its semantic specificity.

However, there are inherent limitations:

  • Synset Ambiguity: The current study only uses the first synset provided by WordNet, which might ignore polysemy (words with multiple meanings, like "Transformers" the movie vs. the electrical device).
  • Static Taxonomy: WordNet is a static human-curated database; it may struggle with modern slang or very specific pop-culture references not yet categorized.

Conclusion & Future Directions

This paper serves as a bridge between symbolic AI (WordNet) and statistical AI (Cosine Similarity). For developers and researchers in the social media space, it highlights that semantic context is more valuable than raw text overlap. Future work could improve this by incorporating Word Embeddings (like Word2Vec or BERT) to handle words not found in traditional lexical databases, further refining the "distance" between human souls in the digital ether.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize Knowledge Graph Embeddings (KGE) instead of manual WordNet hierarchies to calculate node similarity in social networks.
  • Which study first introduced the concept of 'Hypernym Distance' as a numerical feature for NLP tasks, and how has this evolved with the advent of Large Language Models (LLMs)?
  • Explore how semantic similarity measures like WordNet-Cosine are being integrated into Graph Neural Networks (GNNs) for community detection or link prediction.
Contents
Beyond Keywords: Leveraging Lexical Hierarchies for Semantic Node Similarity
1. TL;DR
2. The "Identity Trap" in Social Graphs
3. Methodology: Mapping Language to Space
4. Experimental Validation
4.1. Key Findings:
5. Critical Insight
6. Conclusion & Future Directions