Routing Queries to Experts: Beyond Keyword Matching in Social Search
Routing of Queries in Social Information Retrieval Using Latent and Explicit Semantic Cues
This paper explores Query Routing in Social Information Retrieval (SIR) by identifying subject matter experts within a personal network. It introduces and compares semantic approaches—specifically Latent Dirichlet Allocation (LDA) and a novel link-aware Explicit Semantic Analysis (ESA-Link)—against traditional TF and TF-IDF baselines.
TL;DR
In social information retrieval, finding "who knows what" is more than a string-matching problem. This paper demonstrates that Latent Dirichlet Allocation (LDA) and Explicit Semantic Analysis (ESA-Link)—which leverages Wikipedia's internal link structure—dramatically outperform traditional TF-IDF when routing queries to experts. By understanding the "semantic neighborhood" of a user's expertise, these models ensure that information seekers are connected to the right providers even when their vocabularies don't perfectly align.
The Problem: The "Vocabulary Mismatch" in Expertise Identification
When you ask a question in a social network, your success depends on routing: sending that query to the individual most likely to have the answer. Prior work often treated this as a simple Information Retrieval (IR) task—matching keywords in a query to keywords in a user's profile.
However, this approach fails because:
- Sparsity: Profiles (like a single scientific abstract) are too brief to contain all relevant keywords.
- Semantic Nuance: An expert in "Machine Learning" is likely also an expert in "Neural Networks," even if that specific phrase isn't in their bio.
- Cost of Inquiry: Routing a query to the wrong person creates social "noise" and reduces the likelihood of future engagement.
Methodology: Mining Latent and Explicit Cues
The authors compare two sophisticated ways of looking at "meaning":
1. Latent Dirichlet Allocation (LDA)
Instead of matching words, LDA treats each author's profile as a distribution over topics. If a query matches specific topics (latent variables) that an author frequently writes about, they are considered a high-match candidate.
2. ESA-Link: Explicit Semantics via Wikipedia
While standard ESA maps text to Wikipedia articles, the authors' ESA-Link goes a step further. It uses the link structure of Wikipedia as an Inductive Bias. If a query relates to "Aeronautics" and an author knows about "Fluid Dynamics," the system uses the graph distance between these Wikipedia concepts to calculate relevance.
The similarity formula accounts for the distance between Wikipedia concepts and , effectively rewarding "near-misses" in the semantic space.
The Proof: Experimental Results
Using the Cranfield collection of 1,400 scientific abstracts, the researchers evaluated how well these models could predict the "correct" author for specific queries.
- Semantic Superiority: Both LDA and ESA variants crushed the TF and TF-IDF baselines.
- The Power of Links: ESA-Link provided a noticeable edge over standard ESA, proving that the connectivity of human knowledge (as represented in Wikipedia) is a powerful proxy for expertise.
(a) Shows LDA significantly outperforming TF-IDF at low recall levels—the most critical part of routing.
(c) Highlights that ESA-Link and LDA are the top contenders, performing with comparable accuracy.
Critical Analysis & Conclusion
The takeaway is clear: Context is King. If you are building a social search or "expert finder" system, standard keyword indexing is insufficient.
Limitations:
- The study used a narrow dataset (Aeronautical Engineering), so results might vary in broader or more subjective domains.
- The LDA model used a fixed, small number of topics (10), which might not scale for more diverse knowledge bases.
Future Outlook: While this paper predates the LLM era, its core insight remains relevant. Modern vector databases and RAG (Retrieval-Augmented Generation) systems are the spiritual successors to ESA and LDA. However, the use of explicit graph structures (like Wikipedia links) still offers a level of explainability and structured reasoning that pure "black box" embeddings often lack. Integrating graph-based semantic cues with modern LLMs is likely the next frontier for professional expert-routing systems.
