Personalized Facets: Bridging Social Context and Wikipedia Knowledge for Smarter Search
Personalized Facets for Faceted Search Using Wikipedia Disambiguation and Social Network
This paper introduces a specialized faceted search framework that tackles lexical ambiguity by integrating Wikipedia disambiguation data with social network-derived user profiles. The core method utilizes a personalized graph visualization where vertex sizes are dynamically scaled based on TF-IDF matching between a user's Facebook "Likes" and search results.
TL;DR
Search engines often fail to distinguish between different meanings of the same word. This paper proposes a Personalized Faceted Search system that uses Wikipedia’s disambiguation data to provide semantic context and Facebook user profiles to rank those contexts. The result is a dynamic, graph-based interface where the most relevant "version" of a word is visually emphasized for the specific user.
The "Apple" Problem: Lexical Ambiguity
When you search for "Apple," do you want the fruit, the tech giant, or the record label? Standard search engines rely heavily on popularity, often burying niche but relevant meanings. The authors identify two critical gaps in current information retrieval:
- Lexical Ambiguity: The inability of machines to distinguish senses of a word without context.
- Static Filtering: Search filters (facets) are usually rigid and manually defined, rather than adapting to the user's specific context.
Methodology: The Fusion of Wiki and Social
The authors' approach revolves around three distinct phases to create a "Knowledge Engine" rather than a mere "Search Engine."
1. Data Preparation (The Wiki Backbone)
The system extracts data from Wikipedia dumps. Wikipedia is unique because it contains specific "Disambiguation Pages" designed precisely to solve the ambiguity problem. For example, for "Java," it identifies facets like Places, Computing, and Music.

2. User Profiling (The Facebook Context)
To understand what a user likes, the system connects to the Facebook Graph API. It scrapes the "Likes" of a user, including the name, category, and description of the liked pages. This data is converted into a feature vector using TF-IDF, creating a "Semantic Fingerprint" of the user's interests.
3. Search Visualization (The Heuristic Graph)
The heart of the innovation is the Document Graph Visualization. Instead of a flat list, results are nodes in a graph. The system uses a specific algorithm to calculate the size of each vertex ():
- : The weight of the node based on its match with the user profile.
- : Constant bounds for pixel size.
The algorithm ensures that if a user follows many tech pages, the "Apple Inc." node will appear much larger and more prominent than the "Apple (fruit)" node.
Experimental Insights
The authors evaluated the system using various profiles:
- User "Richard" (Interested in Tech): For the query "Java," the nodes for Programming Language and Software Platform were highlighted.
- User "Tom" (Interested in Nature/Geography): For the same query, Java (island) and Coffee were highlighted.

The study also confirmed that the personalized graph (G2) consumes fewer screen resources and is more efficient for user exploration than a non-personalized graph (G1), as it guides the eye directly to high-probability targets.
Critical Analysis & Future Outlook
Strengths:
- The use of Wikipedia's crowd-sourced disambiguation data provides a very high-quality ground truth for semantic facets.
- The dynamic graph visualization offers a more intuitive UX than traditional sidebar checklists.
Limitations:
- Dependency on Social APIs: The method relies on Facebook's Graph API, which has become significantly more restrictive regarding data privacy since this paper's conception.
- Accuracy Trade-offs: While recall was high, the authors noted that precision can still be improved, as irrelevant documents occasionally surface due to keyword overlap in descriptions.
Conclusion
This work demonstrates that search is not just about indexing documents, but about modeling the user. By linking the world's largest encyclopedia (Wikipedia) with the world's largest social graph (Facebook), the authors provide a compelling blueprint for a search experience that understands intent through association.
