TAP into Social Influence: Quantifying the Invisible Ties in Large Networks
Social influence analysis in large-scale networks
The paper introduces Topical Affinity Propagation (TAP), a novel method for analyzing topic-level social influence in large-scale networks. It utilizes a Topical Factor Graph (TFG) model and an efficient distributed learning algorithm under the Map-Reduce framework, achieving State-of-the-Art performance in quantifying edge-specific influence across different topics.
Executive Summary
TL;DR: This seminal work proposes Topical Affinity Propagation (TAP), a framework that moves beyond "who is important" to "who influences whom on what topic." By combining factor graphs with distributed Map-Reduce learning, it enables the quantitative analysis of influence across millions of nodes, significantly improving expert recommendation tasks.
Background: In the landscape of 2009 (and still relevant today), most social analysis was "topic-blind." A friend might influence your dinner choices but not your programming style. This paper provides the mathematical and computational bridge to distinguish these nuances at scale.
Problem & Motivation: The "Topic-Blind" Limitation
Existing social analysis usually looks at the "macro" level—clustering coefficients or global PageRank. However, these methods ignore two critical realities:
- Multi-aspect Influence: Influence is rarely monolithic. You might follow a mathematician for AI advice but ignore their political views.
- Computational Bottleneck: Standard message-passing algorithms in graphical models (Sum-Product) scale at , making them useless for networks like DBLP or Facebook.
The authors' insight was to treat influence as a "representative selection" problem: for a specific topic, which neighbor "represents" your interest or expertise?
Methodology: Topical Factor Graphs (TFG)
The core of the methodology is the Topical Factor Graph. The authors define a joint distribution over observed nodes () and hidden representative vectors (). Each represents which node has the highest probability of influencing node on topic .
The TFG Architecture
The model relies on three primary functions:
- Node Function (): Captures local topical similarity.
- Edge Function (): Maintains the network topology constraints.
- Global Function (): A specialized constraint ensuring that "representative" nodes are self-consistent (a representative of others must represent itself).

Scaling via Distributed Learning
To solve the problem, the authors derived new update rules for Topical Affinity Propagation. Instead of generic Sum-Product, they use:
- Responsibility (): How much node thinks is a suitable influencer.
- Availability (): How much node thinks it should be an influencer for .
These rules were mapped onto the Map-Reduce framework, allowing the calculation to be split across subgraphs, drastically reducing CPU time.
Experiments & Results: Efficiency and Precision
The authors tested TAP on three massive datasets: DBLP Coauthor, Citation, and a Film-Actor network.
1. Massive Speedup
The distributed TAP implementation showed near-linear speedup. On the Citation dataset (12.7M edges), the basic TAP took over 10 hours, while the distributed version finished in just 39 minutes.

2. Expert Finding (TPRI)
The paper introduced Topic-based PageRank with Influence (TPRI). By replacing uniform transition probabilities with TAP influence scores, they achieved significant gains in Mean Average Precision (MAP) compared to standard PageRank.

Critical Analysis & Conclusion
Takeaway: The TAP framework proved that topic-level features and network structure are not independent—they reinforce each other through propagation. Its ability to quantify influence strengths () provides a much richer feature set for downstream tasks than simple edge weights.
Limitations:
- Topic Independence: The model current assumes topics are independent ( for ), which may not hold in interdisciplinary fields (e.g., "Machine Learning" and "Data Mining").
- Map-Reduce Overhead: For smaller networks, the communication overhead of distributed systems makes basic TAP faster than the distributed version.
Future Outlook: The integration of TAP with modern Graph Neural Networks (GNNs) could potentially allow for representation learning that is inherently aware of topical influence flow, bridging the gap between classical social theory and modern deep learning.
