MEDICI: Synthesizing the "Social" in Social Network Data
MEDICI: A Simple to Use Synthetic Social Network Data Generator
MEDICI is a Java-based synthetic data generator designed to create realistic social network graphs and accompanying user demographics. It integrates the R-MAT algorithm for topology generation, the Louvain method for community detection, and a proprietary seed-propagation mechanism to populate nodes with profiles and "likes."
TL;DR
MEDICI is a comprehensive application that generates synthetic social network datasets, including both the graph structure and rich user metadata (age, gender, religion, "likes"). By combining R-MAT for topology and Louvain for clustering with a specialized Seed-Propagation algorithm, it allows researchers to create realistic, privacy-compliant testbeds for social media analysis.
Motivation: Beyond Just Nodes and Edges
Most synthetic graph generators solve the "structure" problem—they give you a power-law distribution of links. However, they fall short of providing the "meat": who are these users? Why do they follow each other?
The authors of MEDICI recognize that in real networks, homophily (the tendency of individuals to associate with similar others) is the driving force. To simulate this, one cannot simply assign random attributes; there must be a logical flow where a user's community influences their profile.
Methodology: The Core Engine
The generation process follows a rigorous four-step pipeline:
- Topology Generation (R-MAT): Creates the "skeleton" of the network, ensuring it mimics the small-world properties and power-law degree distributions of real social platforms.
- Community Detection (Louvain): Partitions the graph into 10 distinct communities. This creates the "neighborhoods" where specific profiles will dominate.
- Seed Assignment: This is the secret sauce. The algorithm picks "seeds" (influential nodes) that are at least distance-3 apart to prevent overlapping neighborhoods from "polluting" each other's attribute distributions during the first phase of growth.
- Attribute Propagation: Seeds are assigned a strict "Prototype Profile." Their neighbors then "inherit" these traits with a degree of stochastic noise, ensuring the network isn't unnaturally uniform.
Caption: The data propagation mechanism moving from Communities to Seeds, then Neighbors, and finally the rest of the network.
Experiments: Validation through Stability
The researchers tested MEDICI on scales ranging from 1k to 50k nodes. A critical finding was the stability of stochasticity. Even though the process relies on random probability (), the aggregate distributions remain highly consistent across runs.
Caption: Comparison of attribute frequencies across different communities, demonstrating how Community 4 (elderly-biased) differs significantly from the global average.
Performance Stats:
- Graph Generation: Instant for up to 50k nodes.
- Seed Assignment: Most expensive step, taking ~300s for a 50k node graph.
- Stability: Large-scale graphs (50k nodes) showed a deviation of only 0.2% to 0.5% between successive executions, proving the algorithm is robust.
Critical Insight: The "Distance-3" Constraint
Why distance-3? In graph theory, if two seeds are too close, their immediate neighbors overlap. If Seed A is "Political: Left" and Seed B is "Political: Right," a shared neighbor creates a logical conflict in a simple propagation model. By enforcing a buffer, MEDICI allows each "cluster of influence" to establish a clear identity before the final "filling-in" stage handles the boundary nodes.
Conclusion & Future Work
MEDICI lowers the entry barrier for social science researchers. While currently limited to 10 communities and static attributes, the authors plan to incorporate dynamic activity data—simulating how "likes" and posts happen over time. This tool effectively moves synthetic data generation from "random graphs" to "virtual societies."
The MEDICI source code is publicly available on GitHub for researchers looking to extend the framework into multi-modal or temporal social simulations.
