Signed Link Prediction: Solving the Cold-Start Problem via Transfer Learning
5657_Predicting positive and negative links in signed social networks by transfer learning.
This paper introduces a Transfer Learning framework to predict positive and negative (signed) links in social networks, particularly for newly formed "target" networks with sparse labeled data. By leveraging a mature "source" network, the authors propose a combination of generalizable latent topological features and an AdaBoost-based instance weighting algorithm to achieve high-accuracy sign prediction.
TL;DR
In the realm of social networks, a "follow" isn't always a "like." Signed networks (containing both positive and negative links) offer a richer but harder-to-predict map of human interaction. This paper presents a Transfer Learning approach that utilizes mature networks (e.g., an established Epinions graph) to predict link signs in new, data-poor networks. By combining Non-negative Matrix Tri-Factorization (NMTF) for feature extraction and an AdaBoost-like weighting mechanism, the researchers achieved a 40% accuracy boost over traditional methods.
Context: Beyond the "Friendship" Paradigm
Most social network research assumes links are purely positive. However, platforms like Slashdot (Friend vs. Foe) or Wikipedia (Votes for/against) are "signed." The challenge is that in new networks, we rarely have enough labeled data to train a classifier. Furthermore, we cannot rely on the signs of neighboring links (a common assumption in prior work) because they are largely unknown in a new system.
The authors' insight is powerful: while the users change, the underlying social structures (like "the enemy of my friend is my enemy") remain consistent across different platforms.
Methodology: The Architecture of Knowledge Transfer
The authors propose a dual-track strategy: Feature Construction and Instance Weighting.
1. Latent Topological Features via NMTF
Instead of relying on simple metrics like node degree, the authors use Non-negative Matrix Tri-Factorization (NMTF). They decompose the adjacency matrices of both source () and target () networks into a shared latent space ().
This ensures that the features extracted are "generalizable," capturing hidden structural motifs that transcend the specific distribution of a single network.
Figure 1: While explicit features look at local triads (shown above), latent features capture global structural patterns.
2. Adaptive Instance Weighting
Not all data from the source network is helpful; some might be "noise" that leads the model astray. The authors employ an AdaBoost-like algorithm.
- Target instances that are misclassified get increased weights (to force the model to learn them).
- Source instances that are misclassified get decreased weights (to prevent the model from learning patterns that don't fit the target).
Experimental Performance
The researchers tested their approach on three major signed networks: Epinions, Slashdot, and Wikipedia.
SOTA Comparison
The proposed "IW" (Instance Weighting) method consistently outperformed:
- Katz Kernels: A standard matrix-based approach.
- Src/Target Only: Training on only one domain.
- Src+Target: Simply merging datasets without weighting.
Figure 2: Performance across different training data percentages. Note how IW (purple line) maintains the lead even as target data increases.
The Power of Latent Features
A key finding was the effectiveness of NMTF-derived features. In scenarios with very little target data (2%), latent features provided significantly higher accuracy than explicit features like "Betweenness Centrality" or "Edge Embeddedness," proving their superior generalizability.
Figure 3: Comparison of feature types. Latent features (Yellow) consistently outperform classical structural metrics.
Conclusion & Insights
This work demonstrates that for social network analysis, we don't need to reinvent the wheel for every new platform. The "physics" of social interaction—expressed through network topology—is remarkably stable.
Key Takeaways for Practitioners:
- Cold-Start is Solvable: If you are launching a new community, you can seed your moderation or trust algorithms using data from mature, similar platforms.
- Weights Matter: Don't just dump all your data into a training bucket. Use instance weighting to filter out the irrelevant "noise" of the source domain.
- Latent is Stronger than Explicit: Structural patterns captured in latent space (NMTF) are more robust than manually engineered features when crossing domain boundaries.
Limitations: The current approach relies on structural data alone. Future work could integrate content (textual sentiment from comments) with these topological features for even higher precision.
