MGRL: Bridging Social Identities via Multi-Granularity Structural Learning
A Multi-Granularity Representation Learning Framework for User Identification Across Social Networks
The paper introduces MGRL (Multi-Granularity Representation Learning), a framework for User Identity Linkage (UIL) across social networks. It leverages a novel heuristic zippering mechanism and weighted random walks to capture higher-order structural proximity, achieving SOTA performance in identifying corresponding accounts across platforms like Twitter and Foursquare.
Executive Summary
TL;DR: The MGRL framework addresses the "User Identity Linkage" problem—identifying the same natural person across different social platforms (e.g., Twitter and Foursquare)—by treating structural alignment as a multi-granularity representation task. By "zippering" networks through known anchor pairs and using a heuristic weight assigner to guide sampling, it captures the deep, higher-order social circles that simpler models miss.
Academic Positioning: This work bridges Granular Computing (GrC) and Network Representation Learning (NRL), moving beyond 1st and 2nd-order proximity to model global network alignment through guided random walks.
Problem & Motivation: The Gap in Social Circles
Every user leaves a unique structural footprint on a social network defined by their "circle of friends." While user profiles (names, bios) can be faked or hidden, social structures are harder to manipulate. However, existing methods face two major hurdles:
- Structural Complexity: Most models only look at direct friends (1st order) or common neighbors (2nd order), failing to align users who are several hops away from known "anchor points."
- Space Disunity: Learning embeddings for two separate networks often results in different latent spaces, making direct comparison (like cosine distance) unreliable.
The authors' insight is to treat these networks not as separate entities, but as layers of a single, coarser granular structure connected by known identities.
Methodology: Zippering and Heuristic Guidance
The MGRL framework operates through a sophisticated three-step pipeline:
1. The Zippering Operation
Instead of learning two separate embedding models, MGRL "zippers" the Source Network () and Target Network () into a merged graph (). Known Supervisory Anchor Pairs (SAPs) are merged into single nodes. This forces the model to share parameters and iterates toward a unified latent space from the start.

2. Heuristic Weight Assigner
The framework doesn't treat all edges equally. It trains a Heuristic Weight Assigner using the Hadamard product of vertex embeddings: This assigner identifies which edges are "likely" to lead toward hidden anchors. These weights then modify the transition probabilities for a second round of random walks, ensuring the "corpus" of nodes used for training is rich with anchor-oriented context.
3. Multi-Granularity Concatenation
The final representation of a user is a concatenation of their local structural features (from the unweighted graph) and anchor-oriented features (from the weighted graph):
Experiments & Results: Performance Breakthroughs
The model was validated on the Twitter-Foursquare dataset, containing over 1,600 ground-truth anchor pairs.
SOTA Comparison
MGRL consistently outperformed established baselines like IONE and PALE-DeepWalk. Notably, when the training data was scarce (10%), MGRL's Precision@30 was significantly higher, proving that its heuristic guided-sampling is extremely efficient at finding latent connections with minimal supervision.

As shown in the table below, even at high training ratios (90%), MGRL maintains a lead, particularly in the highly competitive Precision@1 and Precision@5 categories, which are critical for real-world identity resolution.
| Method | Precision@1 | Precision@5 | Precision@30 |
|---|---|---|---|
| IONE | 0.2056 | 0.3575 | 0.6012 |
| MGRL (Ours) | 0.2341 | 0.4430 | 0.6456 |
Critical Analysis & Conclusion
Takeaway
MGRL succeeds because it doesn't just look at who you know, but how you are positioned relative to the "anchors" of the network. The "zippering" technique is a powerful inductive bias that resolves the vector space alignment issue that plagues multi-network tasks.
Limitations
- Computational Overhead: Zippering two large-scale networks (e.g., billions of nodes) into one might lead to memory bottlenecks.
- Temporal Dynamics: Social networks evolve. The current model assumes a static snapshot, which might not hold true if a user's friend circle changes significantly between platforms over time.
Future Work: The authors suggest exploring more discriminating structural features beyond simple degrees and triangles, potentially integrating Graph Attention Mechanisms to further refine the heuristic assigner.
