Effects of Missing Data on Social Network Strategy: A Sensitivity Analysis

Effects of missing data in social networks

2005-09-10
Gueorgi Kossinets
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates the sensitivity of social network statistics to missing data by simulating various omission mechanisms on scientific collaboration graphs and bipartite random graphs. It identifies how fundamental metrics like clustering, assortativity, and the giant component size are systematically biased by research design flaws and non-response.

TL;DR

Social network data is rarely "complete," yet our structural conclusions often assume it is. This paper reveals that missing nodes and links don't just add noise—they create systematic biases. Depending on whether you miss an "actor" (an author) or a "context" (a shared paper), your estimates of network clustering and robustness can be wildly over- or underestimated.

Contextualizing the Problem: The Bipartite Reality

In traditional Social Network Analysis (SNA), we often look at one-mode projections—who knows whom. However, real social life is multicontextual. We meet people through specific events, organizations, or collaborations.

The author argues that ignoring this bipartite structure (Actors + Contexts) leads to a failure in understanding the Boundary Specification Problem (BSP). If we arbitrarily cut the boundary of a study, we aren't just losing nodes; we are losing the "cliques" that define the network's social fabric.

Methodology: Simulating "Missingness"

The study utilizes the Los Alamos Condensed Matter (cond-mat) collaboration archive. To understand how "missingness" affects our view of reality, the author applies several simulation models:

  • BSPC (Context Omission): Randomly removing papers.
  • BSPA (Actor Omission): Randomly removing authors.
  • NRE (Non-Response Effect): Removing links that would have been reported by specific individuals.
  • Fixed Choice (Censoring): Simulating surveys that only allow respondents to list "up to X friends."

Model Architecture and Bipartite Table

Key Insights: Why Structural Intuition Fails

1. The Clustering Paradox

One of the most striking findings is the divergent effect on the Clustering Coefficient ():

  • Missing Contexts (BSPC): Actually increases observed clustering. As overlapping cliques are removed, the remaining ones appear more isolated and internally dense.
  • Non-Response (NRE): Decreases clustering by "opening up" triangles into simple pairs or isolated nodes.

2. Assortativity and Robustness

Existing theory (from the Newman/Barabási era) suggests that Assortative networks (where high-degree nodes connect to other high-degree nodes) are more robust to random failures. However, this study finds that in multicontextual networks, this isn't always true. The "core group" of a collaboration network can be surprisingly fragile when authors are missing, despite its assortative nature.

Clustering and Assortativity Sensitivity Figure: The impact of missing data on clustering and connectivity. Note how the error rate () spikes as the fraction of missing data () increases.

3. The Danger of "Fixed Choice" Designs

Many surveys ask: "List your top 5 collaborators." In random graphs, this is fine. But in real-world networks with skewed degree distributions (power laws), this is catastrophic. Because a few "super-connectors" hold the network together, capping nominations at a low number (like 5 or 10) loses the very links that define the network's global connectivity.

Experimental Results on Degree Censoring Figure: Comparison of Fixed Choice effects between real physics collaboration data (a) and random graphs (b). The real network diverges from the "true" mean degree almost immediately.

Conclusion and Deep Insight

The most "ironic" takeaway for any data scientist or sociologist is that multiple errors can hide each other. If your study has both a boundary specification problem (overestimating clustering) and a high non-response rate (underestimating clustering), they might cancel out to produce a "correct-looking" number. However, the underlying measurement error is actually doubled, making any further analysis—such as disease spread modeling or influence mapping—fundamentally flawed.

Future Outlook: As we move toward massive "digital trace" datasets (email logs, social media), we must realize that even "Big Data" is often just a sample bounded by a specific platform context. This paper serves as a rigorous reminder to check the "Redundancy" and "Mixing Patterns" of our data before trusting the metrics.

Find Similar Papers

Try Our Examples

  • Find recent papers that propose statistical imputation or remedial techniques for missing data in bipartite social networks.
  • Which 1983 paper by Laumann et al. established the "boundary specification problem," and how has its definition evolved with the rise of digital trace data?
  • Explore studies investigating the robustness of the "giant component" in networks that exhibit both high clustering and assortative mixing, specifically in the context of disease spreading.
Contents
Effects of Missing Data on Social Network Strategy: A Sensitivity Analysis
1. TL;DR
2. Contextualizing the Problem: The Bipartite Reality
3. Methodology: Simulating "Missingness"
4. Key Insights: Why Structural Intuition Fails
4.1. 1. The Clustering Paradox
4.2. 2. Assortativity and Robustness
4.3. 3. The Danger of "Fixed Choice" Designs
5. Conclusion and Deep Insight