Inf2vec: Revolutionizing Social Influence Modeling via Latent Node Embeddings
Inf2vec: Latent Representation Model for Social Influence Embedding
The paper introduces Inf2vec, a social influence embedding model that represents users in a low-dimensional latent space to capture social influence propagation. Unlike traditional methods that estimate edge-specific probabilities, Inf2vec learns node-level representations by integrating local influence neighborhoods, global user similarities, and individual influence biases.
Executive Summary
TL;DR: Inf2vec is a novel representation learning framework that shifts the paradigm of social influence analysis from estimating local edge probabilities to learning global node embeddings. By combining network topology, propagation history, and latent user interests, it overcomes the "sparsity wall" that plagues traditional models.
Background Positioning: This work bridges the gap between Network Embedding (preserving structure) and Information Diffusion (predicting cascades). It occupies a unique spot in the academic landscape by being one of the first to directly utilize node representations—rather than edge weights—to model the directed, multi-faceted nature of social influence.
The "Sparsity Wall": Why Edge-based Models Fail
Traditional influence models like the Independent Cascade (IC) model rely on having enough data for every single directed edge in a social network. In reality, most edges never see a propagation event. If user A follows user B, but they haven't shared a story yet, traditional models are essentially blind.
The authors' insight is twofold:
- Structural Transitivity: If A influences B and B influences C, A likely has latent influence over C even if no direct "A to C" data exists.
- Latent Preference: Users often perform actions because of shared interests (homophily), not just "viral" pressure. Ignoring user similarity leads to biased influence estimates.
Methodology: The Inf2vec Core
Inf2vec solves the sparsity problem by learning a K-dimensional vector for each user. To handle the directed nature of influence, every user is assigned two vectors:
- Source Vector (): Represents 's ability to influence others.
- Target Vector (): Represents 's susceptibility to being influenced.
1. Generating Multi-faceted Context
The "secret sauce" of Inf2vec is how it defines a user's context. It doesn't just look at direct neighbors; it builds an Influence Propagation Network and samples:
- Local Context: Using random walks with restarts to simulate high-order diffusion.
- Global Context: Sampling users who engaged with the same items, capturing latent interest similarity.
Figure: The process of building influence propagation networks to capture high-order relationships.
2. The Objective Function
The model uses a Word2vec-style skip-gram approach. The probability of user influencing is modeled as: Where is the Influence Ability Bias and is the Conformity Bias. This setup allows the model to distinguish between a "naturally influential" user (like a celebrity) and a specific "influencer-follower" relationship.
Experimental Breakthroughs
The authors tested Inf2vec against heavyweights like DeepWalk, node2vec, and the Embedded Cascade model (Emb-IC).
Performance Comparisons
In the Activation Prediction task (predicting the next person to join a cascade), Inf2vec showed massive gains:
- Digg: Achieved a MAP of 0.2744 (vs. 0.2071 for EM-based models).
- Flickr: Showed significant resilience despite the massive 10M-edge graph.
Table: Comparison of Inf2vec against baselines highlighting its superior precision (P@N) and MAP.
Visualizing Influence
The t-SNE visualizations reveal that Inf2vec successfully clusters frequent influence pairs in the latent space, whereas structure-only models (like node2vec) or similarity-only models (like MF) fail to capture the directed "push and pull" of social dynamics.
Figure: t-SNE visualization comparing various embedding techniques; Inf2vec clusters influent/influenced pairs (marked with symbols) most effectively.
Critical Insight & Future Outlook
Takeaway: Inf2vec proves that "social influence" is not just a property of a link, but a latent property of the nodes themselves. By move from edge-parameters to node-parameters, we gain both efficiency and generalizability.
Limitations: While powerful, Inf2vec currently treats the "Global Similarity" context as static. Future iterations could benefit from Temporal Embeddings to account for how a user's interests—and their social standing—evolve over time. Additionally, integrating Topical Features (NLP) into the embeddings could further refine the "why" behind the influence.
