Routing Queries to Experts: Beyond Keyword Matching in Social Search

Routing of Queries in Social Information Retrieval Using Latent and Explicit Semantic Cues

2016-09-01
Christoph Fuchs, Georg Groh
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores Query Routing in Social Information Retrieval (SIR) by identifying subject matter experts within a personal network. It introduces and compares semantic approaches—specifically Latent Dirichlet Allocation (LDA) and a novel link-aware Explicit Semantic Analysis (ESA-Link)—against traditional TF and TF-IDF baselines.

TL;DR

In social information retrieval, finding "who knows what" is more than a string-matching problem. This paper demonstrates that Latent Dirichlet Allocation (LDA) and Explicit Semantic Analysis (ESA-Link)—which leverages Wikipedia's internal link structure—dramatically outperform traditional TF-IDF when routing queries to experts. By understanding the "semantic neighborhood" of a user's expertise, these models ensure that information seekers are connected to the right providers even when their vocabularies don't perfectly align.

The Problem: The "Vocabulary Mismatch" in Expertise Identification

When you ask a question in a social network, your success depends on routing: sending that query to the individual most likely to have the answer. Prior work often treated this as a simple Information Retrieval (IR) task—matching keywords in a query to keywords in a user's profile.

However, this approach fails because:

  1. Sparsity: Profiles (like a single scientific abstract) are too brief to contain all relevant keywords.
  2. Semantic Nuance: An expert in "Machine Learning" is likely also an expert in "Neural Networks," even if that specific phrase isn't in their bio.
  3. Cost of Inquiry: Routing a query to the wrong person creates social "noise" and reduces the likelihood of future engagement.

Methodology: Mining Latent and Explicit Cues

The authors compare two sophisticated ways of looking at "meaning":

1. Latent Dirichlet Allocation (LDA)

Instead of matching words, LDA treats each author's profile as a distribution over topics. If a query matches specific topics (latent variables) that an author frequently writes about, they are considered a high-match candidate.

2. ESA-Link: Explicit Semantics via Wikipedia

While standard ESA maps text to Wikipedia articles, the authors' ESA-Link goes a step further. It uses the link structure of Wikipedia as an Inductive Bias. If a query relates to "Aeronautics" and an author knows about "Fluid Dynamics," the system uses the graph distance between these Wikipedia concepts to calculate relevance.

Formula for ESA-Link Similarity The similarity formula accounts for the distance between Wikipedia concepts and , effectively rewarding "near-misses" in the semantic space.

The Proof: Experimental Results

Using the Cranfield collection of 1,400 scientific abstracts, the researchers evaluated how well these models could predict the "correct" author for specific queries.

  • Semantic Superiority: Both LDA and ESA variants crushed the TF and TF-IDF baselines.
  • The Power of Links: ESA-Link provided a noticeable edge over standard ESA, proving that the connectivity of human knowledge (as represented in Wikipedia) is a powerful proxy for expertise.

Performance Comparison Graph (a) Shows LDA significantly outperforming TF-IDF at low recall levels—the most critical part of routing.

ESA Results Comparison (c) Highlights that ESA-Link and LDA are the top contenders, performing with comparable accuracy.

Critical Analysis & Conclusion

The takeaway is clear: Context is King. If you are building a social search or "expert finder" system, standard keyword indexing is insufficient.

Limitations:

  • The study used a narrow dataset (Aeronautical Engineering), so results might vary in broader or more subjective domains.
  • The LDA model used a fixed, small number of topics (10), which might not scale for more diverse knowledge bases.

Future Outlook: While this paper predates the LLM era, its core insight remains relevant. Modern vector databases and RAG (Retrieval-Augmented Generation) systems are the spiritual successors to ESA and LDA. However, the use of explicit graph structures (like Wikipedia links) still offers a level of explainability and structured reasoning that pure "black box" embeddings often lack. Integrating graph-based semantic cues with modern LLMs is likely the next frontier for professional expert-routing systems.

Find Similar Papers

Try Our Examples

  • What are the most recent SOTA methods for expert finding in social networks that combine Knowledge Graphs with Transformer-based embeddings?
  • Examine the foundational papers for Explicit Semantic Analysis (ESA) and how newer models like D-ESA or Neural ESA have improved upon the original concept-mapping logic.
  • How has the specific problem of Query Routing in Social Information Retrieval been adapted for large-scale decentralized platforms like Mastodon or modern enterprise Slack environments?
Contents
Routing Queries to Experts: Beyond Keyword Matching in Social Search
1. TL;DR
2. The Problem: The "Vocabulary Mismatch" in Expertise Identification
3. Methodology: Mining Latent and Explicit Cues
3.1. 1. Latent Dirichlet Allocation (LDA)
3.2. 2. ESA-Link: Explicit Semantics via Wikipedia
4. The Proof: Experimental Results
5. Critical Analysis & Conclusion