The Power of Bias: Why Random Sampling Fails Social Influence Estimation
Estimating influence of social media users from sampled social networks
This paper investigates how various node sampling techniques affect the accuracy of user influence estimation in social networks. By comparing indices like PageRank and Degree Centrality across four real-world datasets (Twitter, Facebook, APS Journals), the authors find that biased sampling methods (like SEC) significantly outperform random sampling in identifying top influencers.
TL;DR
In the world of Social Network Analysis (SNA), more data isn't always better—but the right data is everything. This paper demonstrates that randomly sampling a social network makes influence indices (like PageRank) almost useless. However, by using biased sampling methods like Sample Edge Count (SEC) or Breadth-First Search (BFS), we can identify the top 1% of influencers with nearly the same accuracy as a complete network analysis, even when looking at only a fraction of the nodes.
Background: The Incompleteness Problem
Identifying "superspreaders" is the holy grail of viral marketing. We rely on centrality measures (Degree, Betweenness, etc.) to find them, but there is a catch: we rarely have the "full" graph of Facebook or Twitter. We usually work with sampled snapshots.
The authors tackle a critical question: How much does the sampling method distort our view of who is actually influential?
The "Overlap 1%" Metric: Measuring Reality
Unlike theoretical studies that use simulated models (like SIR or IC), this study uses ground truth cascade data. They define actual influence by:
- Twitter: The number of unique users who retweeted a post.
- Facebook: The number of posts on a user's wall.
- APS Journals: The total citations of an author's papers.
They then use Overlap 1%—the ratio of top influencers correctly identified in the sampled network compared to the actual top 1% of the full dataset.
Methodology: Sampling Strategies
The researchers compared four approaches to "gathering" a network:
- Random Sampling: Picking nodes like a lottery.
- BFS/DFS: Traversal methods that explore local neighborhoods.
- SEC (Sample Edge Count): A greedy approach that prioritizes nodes with the most links to the already-sampled set.
Table 1: Statistics of the four diverse datasets used to validate the findings.
Deep Dive into the Results
The findings offer a stark contrast between "fair" sampling and "useful" sampling.
1. The Failure of Randomness
As shown in the charts, as the sample size decreases, Random Sampling's ability to find influencers collapses. At a 10% sample, the overlap is just 0.05. This suggests that if you randomly crawl 10% of a network, you lose the structural "high ground" needed for PageRank or k-core to function.
(Figure 1: The rapid decay of Overlap 1% as random sample size decreases across datasets.)
2. The Biased Advantage
Curiously, methods that are "biased" toward high-degree nodes (SEC, BFS) stay resilient. Even at tiny sample sizes, SEC maintains a high Overlap 1%. Why? Because social network influence is often concentrated in the "hubs." Biased sampling acts as a filter that preserves these critical nodes and their connections.
(Figure 2: Performance of SEC sampling—note how the curves stay high even at low sample sizes.)
Critical Insights: Degree Centrality vs. Global Indices
A surprising takeaway is that Degree Centrality (a local measure) often performs as well as, or better than, PageRank (a global measure) in sampled networks. Global measures like PageRank require the "whole picture" to be accurate; when segments of the graph are missing, their calculations become skewed. Degree Centrality, being local, is more robust to the absence of distant nodes.
Conclusion and Future Outlook
This paper changes the narrative on network data collection. If your goal is to identify leaders or influencers:
- Avoid random node selection. It breaks the topology that influence indices rely on.
- Embrace traversal-based sampling. Methods like SEC effectively "pre-filter" the network for the most important nodes.
Limitations: The authors note that in specific types of networks (like the APS Citation network), even biased sampling struggles if influencers are widely distributed across isolated clusters. Future work will need to focus on "topology-aware" sampling that can bridge these gaps.
