Effects of Missing Data on Social Network Strategy: A Sensitivity Analysis
Effects of missing data in social networks
This paper investigates the sensitivity of social network statistics to missing data by simulating various omission mechanisms on scientific collaboration graphs and bipartite random graphs. It identifies how fundamental metrics like clustering, assortativity, and the giant component size are systematically biased by research design flaws and non-response.
TL;DR
Social network data is rarely "complete," yet our structural conclusions often assume it is. This paper reveals that missing nodes and links don't just add noise—they create systematic biases. Depending on whether you miss an "actor" (an author) or a "context" (a shared paper), your estimates of network clustering and robustness can be wildly over- or underestimated.
Contextualizing the Problem: The Bipartite Reality
In traditional Social Network Analysis (SNA), we often look at one-mode projections—who knows whom. However, real social life is multicontextual. We meet people through specific events, organizations, or collaborations.
The author argues that ignoring this bipartite structure (Actors + Contexts) leads to a failure in understanding the Boundary Specification Problem (BSP). If we arbitrarily cut the boundary of a study, we aren't just losing nodes; we are losing the "cliques" that define the network's social fabric.
Methodology: Simulating "Missingness"
The study utilizes the Los Alamos Condensed Matter (cond-mat) collaboration archive. To understand how "missingness" affects our view of reality, the author applies several simulation models:
- BSPC (Context Omission): Randomly removing papers.
- BSPA (Actor Omission): Randomly removing authors.
- NRE (Non-Response Effect): Removing links that would have been reported by specific individuals.
- Fixed Choice (Censoring): Simulating surveys that only allow respondents to list "up to X friends."

Key Insights: Why Structural Intuition Fails
1. The Clustering Paradox
One of the most striking findings is the divergent effect on the Clustering Coefficient ():
- Missing Contexts (BSPC): Actually increases observed clustering. As overlapping cliques are removed, the remaining ones appear more isolated and internally dense.
- Non-Response (NRE): Decreases clustering by "opening up" triangles into simple pairs or isolated nodes.
2. Assortativity and Robustness
Existing theory (from the Newman/Barabási era) suggests that Assortative networks (where high-degree nodes connect to other high-degree nodes) are more robust to random failures. However, this study finds that in multicontextual networks, this isn't always true. The "core group" of a collaboration network can be surprisingly fragile when authors are missing, despite its assortative nature.
Figure: The impact of missing data on clustering and connectivity. Note how the error rate () spikes as the fraction of missing data () increases.
3. The Danger of "Fixed Choice" Designs
Many surveys ask: "List your top 5 collaborators." In random graphs, this is fine. But in real-world networks with skewed degree distributions (power laws), this is catastrophic. Because a few "super-connectors" hold the network together, capping nominations at a low number (like 5 or 10) loses the very links that define the network's global connectivity.
Figure: Comparison of Fixed Choice effects between real physics collaboration data (a) and random graphs (b). The real network diverges from the "true" mean degree almost immediately.
Conclusion and Deep Insight
The most "ironic" takeaway for any data scientist or sociologist is that multiple errors can hide each other. If your study has both a boundary specification problem (overestimating clustering) and a high non-response rate (underestimating clustering), they might cancel out to produce a "correct-looking" number. However, the underlying measurement error is actually doubled, making any further analysis—such as disease spread modeling or influence mapping—fundamentally flawed.
Future Outlook: As we move toward massive "digital trace" datasets (email logs, social media), we must realize that even "Big Data" is often just a sample bounded by a specific platform context. This paper serves as a rigorous reminder to check the "Redundancy" and "Mixing Patterns" of our data before trusting the metrics.
