MCULK: Bridging the Identity Gap Across Multiple Social Platforms

User Profile Linkage Across Multiple Social Platforms

2020-01-01
Manman Wang, Wei Chen, Jiajie Xu, Pengpeng Zhao, Lei Zhao
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces MCULK, a multi-platform User Profile Linkage (UPL) framework designed to unify identities across three or more social networks. Unlike pairwise methods, MCULK utilizes a similarity graph construction combined with localized clustering to achieve state-of-the-art performance in resolving "platform uncertainty."

TL;DR

Standard User Profile Linkage (UPL) often breaks down when moving beyond two platforms. MCULK solves this by shifting from binary matching to a similarity graph clustering approach. By optimizing the process with Locality Sensitive Hashing (LSH) and a refined split-and-merge clustering logic, it achieves SOTA accuracy while being orders of magnitude faster than traditional iterative methods.

Problem & Motivation: The "Platform Uncertainty" Trap

In the modern digital landscape, a single user is rarely confined to one platform. While we might have a Twitter for news, an Instagram for photos, and a LinkedIn for work, current UPL research still largely obsesses over pairwise linkage (Platform A vs. Platform B).

When you try to scale pairwise results to m platforms, you encounter two major walls:

  1. High Computational Cost: Comparing every pair across platforms leads to complexity.
  2. Incorrect Linkage Chains: If A matches B, and B matches C, a single error in one pairwise link can cause a "chain reaction," incorrectly merging unrelated users into a massive, erroneous cluster.

The authors identify "Platform Uncertainty"—the fact that users exist on an unpredictable subset of services—as the primary obstacle that simple pairwise logic cannot handle.

Methodology: From Similarity Graphs to Consistent Clusters

MCULK (Multi-platform Cluster User Linkage) approaches the problem in two distinct phases:

1. Efficient Similarity Graph Generation

To avoid the quadratic explosion, the model uses MinHashLSH (Locality Sensitive Hashing). Profiles are "bucketed" based on the q-gram similarity of their usernames and display names. Only profiles within the same bucket are passed to a pre-trained Logistic Regression classifier. This produces a weighted undirected graph where nodes are profiles and edges represent the probability of belonging to the same human.

System Architecture

2. Adaptive Clustering Logic

The "magic" happens in how the similarity graph is partitioned:

  • Initial Clustering: Uses connected components but strictly enforces source-consistency (a cluster cannot contain two profiles from the same platform).
  • Splitting: If a cluster’s internal similarity is too low (below threshold ), it is broken apart to ensure high precision.
  • Merging: It creates "representatives" for incomplete clusters and attempts to merge them using index-based lookups, ensuring that we don't miss links between platforms that were not directly compared in the LSH phase.

Experiments & Results: Speed Meets Precision

The researchers tested MCULK on two massive datasets (DS1: Google+, Twitter, Instagram; DS2: A 100k+ profile dataset across four sources).

SOTA Comparison

In terms of F1 Score, MCULK consistently outperformed established methods like OPL and UISN-UD. More importantly, it solved the "complete cluster" problem where users on all platforms must be linked correctly.

Performance Comparison

Efficiency Breakthrough

The performance trade-off is virtually non-existent. On the DS2 dataset, MCULK completed the linkage task in 0.092 seconds, while the competition took up to 49.9 seconds. This represents a massive leap in scalability for production environments.

Efficiency Table

Critical Insight & Conclusion

The core takeaway from this work is that UPL is a clustering problem, not a sequence of classification problems. By treating the data as a graph, MCULK maintains a global view of identity that pairwise methods lack.

Limitations: The model relies heavily on username and name similarity. If a user adopts completely distinct personas (e.g., "TechWizard" on Twitter vs. "JohnDoe" on LinkedIn), the q-gram blocking might miss them.

Future Outlook: The authors suggest incorporating User Generated Content (UGC) and Friendship Networks as the next frontier. Imagine a model that links you not because your name matches, but because your writing style or your social circle is identical across platforms. This work provides the architectural foundation for that reality.

Find Similar Papers

Try Our Examples

  • Find recent papers on user identity linkage (UIL) that utilize Graph Neural Networks (GNNs) for better structural similarity representation across large-scale social networks.
  • What are the seminal papers on "blocked clustering" or "incremental entity resolution" that influenced the development of the FAMER framework mentioned in this study?
  • Explore how multi-platform user linkage techniques are currently being applied to enhance cross-domain recommendation systems or to detect coordinated inauthentic behavior (CIB).
Contents
MCULK: Bridging the Identity Gap Across Multiple Social Platforms
1. TL;DR
2. Problem & Motivation: The "Platform Uncertainty" Trap
3. Methodology: From Similarity Graphs to Consistent Clusters
3.1. 1. Efficient Similarity Graph Generation
3.2. 2. Adaptive Clustering Logic
4. Experiments & Results: Speed Meets Precision
4.1. SOTA Comparison
4.2. Efficiency Breakthrough
5. Critical Insight & Conclusion