Personalized Facets: Bridging Social Context and Wikipedia Knowledge for Smarter Search

Personalized Facets for Faceted Search Using Wikipedia Disambiguation and Social Network

2016-01-01
Hong Son Nguyen, Hong Phuc Pham, Trong Hai Duong, Thi Phuong Trang Nguyen, Triet Huynh Minh Le
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a specialized faceted search framework that tackles lexical ambiguity by integrating Wikipedia disambiguation data with social network-derived user profiles. The core method utilizes a personalized graph visualization where vertex sizes are dynamically scaled based on TF-IDF matching between a user's Facebook "Likes" and search results.

TL;DR

Search engines often fail to distinguish between different meanings of the same word. This paper proposes a Personalized Faceted Search system that uses Wikipedia’s disambiguation data to provide semantic context and Facebook user profiles to rank those contexts. The result is a dynamic, graph-based interface where the most relevant "version" of a word is visually emphasized for the specific user.

The "Apple" Problem: Lexical Ambiguity

When you search for "Apple," do you want the fruit, the tech giant, or the record label? Standard search engines rely heavily on popularity, often burying niche but relevant meanings. The authors identify two critical gaps in current information retrieval:

  1. Lexical Ambiguity: The inability of machines to distinguish senses of a word without context.
  2. Static Filtering: Search filters (facets) are usually rigid and manually defined, rather than adapting to the user's specific context.

Methodology: The Fusion of Wiki and Social

The authors' approach revolves around three distinct phases to create a "Knowledge Engine" rather than a mere "Search Engine."

1. Data Preparation (The Wiki Backbone)

The system extracts data from Wikipedia dumps. Wikipedia is unique because it contains specific "Disambiguation Pages" designed precisely to solve the ambiguity problem. For example, for "Java," it identifies facets like Places, Computing, and Music.

Document features extraction process

2. User Profiling (The Facebook Context)

To understand what a user likes, the system connects to the Facebook Graph API. It scrapes the "Likes" of a user, including the name, category, and description of the liked pages. This data is converted into a feature vector using TF-IDF, creating a "Semantic Fingerprint" of the user's interests.

3. Search Visualization (The Heuristic Graph)

The heart of the innovation is the Document Graph Visualization. Instead of a flat list, results are nodes in a graph. The system uses a specific algorithm to calculate the size of each vertex ():

  • : The weight of the node based on its match with the user profile.
  • : Constant bounds for pixel size.

The algorithm ensures that if a user follows many tech pages, the "Apple Inc." node will appear much larger and more prominent than the "Apple (fruit)" node.

Experimental Insights

The authors evaluated the system using various profiles:

  • User "Richard" (Interested in Tech): For the query "Java," the nodes for Programming Language and Software Platform were highlighted.
  • User "Tom" (Interested in Nature/Geography): For the same query, Java (island) and Coffee were highlighted.

Experimental evaluation of vertex size

The study also confirmed that the personalized graph (G2) consumes fewer screen resources and is more efficient for user exploration than a non-personalized graph (G1), as it guides the eye directly to high-probability targets.

Critical Analysis & Future Outlook

Strengths:

  • The use of Wikipedia's crowd-sourced disambiguation data provides a very high-quality ground truth for semantic facets.
  • The dynamic graph visualization offers a more intuitive UX than traditional sidebar checklists.

Limitations:

  • Dependency on Social APIs: The method relies on Facebook's Graph API, which has become significantly more restrictive regarding data privacy since this paper's conception.
  • Accuracy Trade-offs: While recall was high, the authors noted that precision can still be improved, as irrelevant documents occasionally surface due to keyword overlap in descriptions.

Conclusion

This work demonstrates that search is not just about indexing documents, but about modeling the user. By linking the world's largest encyclopedia (Wikipedia) with the world's largest social graph (Facebook), the authors provide a compelling blueprint for a search experience that understands intent through association.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize Knowledge Graphs and Graph Neural Networks for personalized faceted search in e-commerce.
  • What are the current SOTA methods for automatic facet generation in unstructured document collections without using Wikipedia as a backbone?
  • How has the integration of social media profiles for search personalization evolved in the era of privacy-preserving machine learning and GDPR?
Contents
Personalized Facets: Bridging Social Context and Wikipedia Knowledge for Smarter Search
1. TL;DR
2. The "Apple" Problem: Lexical Ambiguity
3. Methodology: The Fusion of Wiki and Social
3.1. 1. Data Preparation (The Wiki Backbone)
3.2. 2. User Profiling (The Facebook Context)
3.3. 3. Search Visualization (The Heuristic Graph)
4. Experimental Insights
5. Critical Analysis & Future Outlook
6. Conclusion