TSCEFM: Bridging Topical Authority and Contextual Relevance in Social Expert Retrieval
A Topic-Specific Contextual Expert Finding Method in Social Network
The paper introduces TSCEFM, a Topic-Specific Contextual Expert Finding Method designed for social networks. It integrates a Topic-Aware Model (TAM) using LDA and HITS with a Context-Aware Model (CAM) evaluating social, temporal, and spatial factors to rank experts via an SVM scoring function.
TL;DR
Finding an "expert" isn't just about who knows the most; it's about who is relevant right now and here. This paper presents TSCEFM, a framework that combines Latent Dirichlet Allocation (LDA) and HITS algorithms with a multi-factor Context-Aware Model. By leveraging SVM to learn a scoring function from social, temporal, and spatial features, it achieves a ~13% boost in ranking accuracy over existing SOTA models.
Problem & Motivation: The Noise in the Network
In the era of social media, everyone is a "content creator," but few are true authorities. Most existing expert finding systems suffer from two main flaws:
- Topic Irrelevance: They analyze the entire social graph, including users and interactions that have nothing to do with the specific query (e.g., a JAVA expert being ranked based on their popularity in cooking).
- Context Blindness: Expertise is often contextual. You might need an expert who is currently active (Time), nearby (Location), or within your extended social circle (Social Relation).
The authors argue that a truly "practical" expert must sit at the intersection of Topical Authority and Contextual Accessibility.
Methodology: The TSCFM Framework
The core of the paper is the Topic-Specific Contextual Feature Model (TSCFM), which breaks down the problem into two parallel tracks:
1. Topic-Aware Model (TAM)
The TAM focuses on "what" the user knows. It uses:
- LDA (Latent Dirichlet Allocation): To build a User Topic Feature Matrix (UTFM), determining the probability that a user belongs to a specific topic.
- Modified HITS: Unlike the traditional HITS algorithm, the authors introduce a Topic-Relevant Base Set (TRBS). They only expand the graph using nodes that have topic-relevant links, effectively filtering out the noise of the general social network.
- Behavior Weighting: It distinguishes between "reading" and "commenting," giving more weight to active engagement ().
2. Context-Aware Model (CAM)
The CAM handles the "when" and "where":
- Social Relation (): Calculated via interaction frequency.
- Location Factor (): Uses a Kernel Density Estimation (KDE) to map the geographical relevance of a user.
- Temporal Factor (): Represents user activity as time vectors and uses cosine similarity to find experts active at the same time as the seeker.

3. The Scoring Function
Rather than manually weighting these features, the authors use an SVM (Support Vector Machine) to learn the scoring function , which takes the combined topical and contextual vector as input to output an authoritative rank.
Experiments & Results
The model was validated on two distinct datasets: CiteSeer (academic citations) and Twitter (social interactions).
Key Findings:
- Superior Accuracy: TSCEFM consistently outperformed baseline models like SAR, TWG, and traditional HITS across both MAP and NDCG metrics.
- The Weight of Topics: Ablation studies revealed that while both factors are necessary, the topical factor () influences accuracy more significantly than the contextual factor ().
- The Power of Fusion: Combining both factors (TSCFM) resulted in a 16.82% improvement in MAP compared to using TAM alone.
Fig: MAP and NDCG results showing the stability and performance of TSCEFM on Twitter and CiteSeer datasets.
Critical Analysis & Conclusion
Takeaway
TSCEFM proves that "expertness" is a multi-dimensional construct. By filtering the social graph into a "Topic-Relevant Base Set" before running authority algorithms (HITS), the researchers successfully eliminated the bias of topic-irrelevant popular users.
Limitations & Future Work
The current model assumes a relatively static social network during the inference phase. The authors acknowledge that social networks are highly dynamic—relationships and topic trends shift rapidly. Future research will likely need to incorporate streaming data or dynamic graph embeddings to maintain accuracy over time.
Conclusion
For developers building recommendation engines or Q&A platforms, this paper provides a robust blueprint: don't just look for the most "popular" nodes; look for the most "relevant" nodes within the specific contextual slice of the user's needs.
