Beyond Keywords: Leveraging Network Propagation for High-Precision Expert Finding

Expert Finding in a Social Network

2007-08-02
Jing Zhang, Jie Tang, Juan-Zi Li
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a propagation-based framework for expert finding in large-scale social networks by integrating local person information with network-level relationships. The method leverages a two-step process—initialization via probabilistic IR and iterative propagation—to achieve state-of-the-art performance on academic researcher discovery.

TL;DR

Finding the right expert in a sea of millions is a daunting task. This paper from Tsinghua University moves beyond simple keyword matching by treating expert discovery as a propagation problem on a social graph. By combining a researcher's local profile with their collaborative network (coauthor relationships), the authors achieved a massive boost in precision, proving that "who you know" is as important as "what you've written."

The "Topic Drift" Problem in Expert Discovery

In the early days of social network mining, researchers typically fell into two camps:

  1. Content-based Search: Using Information Retrieval (IR) to find people whose papers match a query.
  2. Link-based Ranking: Using PageRank or HITS to find influential nodes.

The problem? Content-based search ignores the rich social context (the validation of peers), while PageRank suffers from topic drift—where globally popular individuals dominate the results even if their specific expertise in a niche topic is low. Furthermore, existing benchmarks were too small to reflect the complexity of a network containing nearly half a million nodes.

The Core Innovation: A Two-Stage Unified Framework

The authors argue that expertise should be calculated through a unified approach. Their methodology follows a elegant "Initialize then Propagate" workflow.

1. Initialization (Local Context)

First, the system aggregates all local information for a person—profiles, contact info, and publications—into a single virtual document. A probabilistic IR model calculates an initial "Expertise Score" () based on how well this document matches the search query.

2. Propagation (Network Context)

This is where the magic happens. The initial scores are propagated through the network using a weighted iterative process:

This formula essentially says: Your expertise at step depends on your current score plus the scores of your neighbors, weighted by the strength of your relationship (). If you collaborate with top-tier experts, your own "stock" rises.

Model Architecture and Network Snippet Figure 1: A snippet of the academic network showing complex relationships like 'supervised_by' and 'coauthor' which drive the propagation process.

Experimental Results: Relationship Matters

The researchers tested their approach on an academic dataset of 448,289 persons and 2.4 million coauthor relationships. The results were conclusive:

MetricBaseline (IR Only)Propagation ApproachImprovement
Precision@546.1561.54+33.3%
MAP9.7311.03+13.3%
bpref13.6416.11+18.1%

Performance Comparison Table 1: Comparing the baseline against the propagation approach across various precision metrics.

The significant jump in Precision@5 suggests that for the top-ranked candidates, the network structure provides a critical "sanity check" that local text alone cannot offer.

Critical Insight & Outlook

This work serves as a precursor to modern Graph Neural Networks (GNNs). It highlights a fundamental truth in social computing: Expertise is contagious.

While the paper focuses on academic coauthorship, the implications are broad. The same logic applies to:

  • Enterprise Search: Finding the right engineer based on their interaction in Slack or GitHub.
  • Knowledge Management: Identifying internal subject matter experts who may not have "expert" in their job title but are central to the organization's collaborative graph.

Limitations: The model currently treats all relationships within a type equally (uniform weights). Future iterations could benefit from Temporal Dynamics—weighting a coauthorship from 2024 higher than one from 1994—to identify current, rather than historical, experts.

Conclusion

By bridge the gap between IR and Link Analysis, Zhang et al. provided a robust, scalable blueprint for expertise discovery. It proves that in the digital age, your value is defined not just by what you store in your local "profile," but by your position within the global knowledge network.

Find Similar Papers

Try Our Examples

  • Find recent papers on graph neural networks (GNNs) for expert finding that utilize both textual features and network topology.
  • Which paper introduced the fundamental Theory of Propagation used in this study, and how has Belief Propagation evolved for social network mining?
  • Search for research applying this propagation-based expertise model to other domains like LinkedIn for industry expert discovery or GitHub for developer skill assessment.
Contents
Beyond Keywords: Leveraging Network Propagation for High-Precision Expert Finding
1. TL;DR
2. The "Topic Drift" Problem in Expert Discovery
3. The Core Innovation: A Two-Stage Unified Framework
3.1. 1. Initialization (Local Context)
3.2. 2. Propagation (Network Context)
4. Experimental Results: Relationship Matters
5. Critical Insight & Outlook
6. Conclusion