ITERGM: Mastering the "Invisible" in Evolving Social Networks

Imputation of missing links and attributes in longitudinal social surveys

2013-10-18
Vladimir Ouzienko, Zoran Obradovic
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces ITERGM (Iterative Temporal Exponential Random Graph Model), a unified framework for the joint imputation of missing links and actor attributes in longitudinal social networks. By leveraging the interdependence between network topology and behavioral attributes over time, it achieves state-of-the-art accuracy on both synthetic and real-world datasets like 'Delinquency' and 'Teenagers'.

TL;DR

Social scientists often face "non-respondent" gaps in surveys where both a person's connections and their traits (e.g., smoking habits) are missing. ITERGM solves this by treating link formation and attribute changes as a single, co-evolving process. By iteratively predicting one based on the other across time, it recovers missing social structures with precision that traditional static methods cannot match.

Background: The High Cost of Silence

In Social Network Analysis (SNA), a missing respondent isn't just a missing row in a table; it's a hole in the "fabric" of the group. If a popular student doesn't fill out a survey, we miss their outgoing links, leading to an underestimation of the group's density and reciprocity.

The core insight of this paper is Homophily Selection: people choose friends similar to themselves, and friends influence each other's behaviors. Previous methods like Preferential Attachment or DynaMMo saw only half the picture—either the graph or the data sequence. ITERGM sees both.

Methodology: The Power of Iteration

ITERGM transforms a prediction model (etERGM) into an imputation engine using an EM-like cycle.

1. Initialization

The model starts by filling gaps using baseline methods (DynaMMo for attributes and a density-based "best guess" for links) to create a "temporary" complete dataset.

2. The Feedback Loop

The algorithm then enters an interlocking training phase:

  • Step A: Train an attribute model. Sample new attributes based on the current (possibly noisy) network structure.
  • Step B: Train a link model. Sample new links based on the newly updated attributes.
  • Step C: Repeat until convergence (usually 3-4 iterations).

3. Architecture & Logic

The model relies on log-linear statistics (sufficient statistics) that measure things like Reciprocity and Transitivity.

Figure 1: Comparison of Link Imputation Accuracy Above: Note how ITERGM (top line) maintains superior AUC even as the number of actors increases compared to simple reconstruction or preferential attachment.

Experiments: Real-World Classrooms

The authors tested ITERGM on two famous longitudinal datasets:

  1. Delinquency: Dutch students and their delinquent behavior over 4 waves.
  2. Teenagers: Alcohol consumption and friendship patterns over 3 waves.

They simulated four types of "Missingness," including Missing Not At Random (MNAR)—for instance, assuming that students with higher delinquency scores are less likely to respond.

Key Performance Metrics:

  • Link Recovery: ITERGM achieved AUC scores significantly higher than baselines, often staying above 0.80 even with 20-40% of nodes missing.
  • Attribute Precision: By using the network as a "prior," attribute MSE was lower than standard averaging techniques.
  • Scalability: The runtime scales linearly with the number of survey waves and quadratically with the number of actors (), which is the theoretical minimum for considering all possible pair-wise relationships.

Table 2: Results on Delinquency Dataset The results show that even under MNAR-outdegree conditions, ITERGM provides the most stable and accurate link imputation.

Critical Insight & Limitations

ITERGM is a significant step forward because it respects the Temporal Inductive Bias—the idea that today’s friendship is likely a function of yesterday’s friendship and today’s mutual interests.

Limitations:

  • The Cold Start Problem: Currently, the model cannot impute the very first time step (Wave 1) because it lacks a "previous" wave to learn the transition.
  • Model Degeneracy: Like many ERGMs, the link prediction can sometimes become "degenerate" (converging to empty or full graphs) if the statistics are not carefully tuned, especially when reversing the time sequence.

Conclusion

This work provides a robust toolkit for social scientists. Instead of discarding incomplete surveys, researchers can use ITERGM to reconstruct the "missing links" of human interaction with statistical confidence, ensuring that their later findings on social influence and group dynamics are based on a complete and accurate picture.

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend Exponential Random Graph Models (ERGMs) for joint link-attribute inference in large-scale dynamic networks.
  • Which original study first proposed the Extended Temporal Exponential Random Graph Model (etERGM), and what were its primary limitations regarding missing data?
  • Are there any studies applying iterative EM-based imputation techniques to multi-layer or multiplex social networks where nodes have multiple types of relationships?
Contents
ITERGM: Mastering the "Invisible" in Evolving Social Networks
1. TL;DR
2. Background: The High Cost of Silence
3. Methodology: The Power of Iteration
3.1. 1. Initialization
3.2. 2. The Feedback Loop
3.3. 3. Architecture & Logic
4. Experiments: Real-World Classrooms
4.1. Key Performance Metrics:
5. Critical Insight & Limitations
6. Conclusion