Bridging the Digital Divide: Recursive Identity Discovery Across Social Networks
Discovering missing me edges across social networks q
This paper introduces a specialized method for discovering missing "me edges"—links between accounts belonging to the same individual—across disparate social networks. It proposes a recursive common-neighbor approach combined with string similarity that achieves high precision in real-life Social Internetworking Scenarios (SIS).
TL;DR
As users fragment their digital lives across platforms like Twitter, LinkedIn, and Instagram, the "me edge"—a formal link between two accounts owned by the same person—is frequently missing. This paper proposes a recursive, node-similarity algorithm that overcomes the limitations of traditional link prediction. By combining username string analysis with a specialized "bridge-assortativity" heuristic, the authors can identify missing identity links with over 90% precision, effectively unifying fragmented social graphs.
The Identity Fragmentation Problem
In the era of Social Internetworking Scenarios (SIS), we are no longer defined by a single social graph but by a "constellation" of networks. While standards like FOAF (Friend-Of-A-Friend) and XFN allow users to explicitly declare "this is me" on another site, most users neglect to do so.
This creates a massive "missing link" problem. Traditional Social Network Analysis (SNA) tools are designed for single-network silos. When you try to find a user across two different networks, standard Common-Neighbor algorithms fail because the neighbors themselves are "anonymous" relative to the other network. It's a recursive cold-start problem: you can't find the user's identity without knowing their friends' identities, but you can't know the friends' identities without a unified graph.
Methodology: Recursive Similarity & Public Figure Filtering
The authors' solution is an elegant recursive operator . Instead of a static check, the algorithm propagates similarity through the network layers.
1. The Similarity Operator
The similarity of two accounts is defined by:
- String Similarity: Using metrics like QGrams or Levenshtein to compare usernames (e.g.,
johnsmithvsdj_smith). - Weighted Neighbor Overlap: A recursive step where the similarity of friends' accounts contributes to the root pair's score.
2. The "Public Figure" Problem
A major pitfall in graph analysis is the "Obama Effect." If two people both follow a global celebrity, it doesn't mean they are the same person; it just means they are both on the internet. The paper introduces a scaling coefficient that diminishes the weight of accounts with extremely high indegree, effectively filtering out the noise of mass-followings.
Figure 1: The Bridge-Assortativity logic, finding candidates by traversing existing me-edges of neighbors.
Experiments: Why Traditional SOTA Fails
The authors pitted their method against established indices like the Jaccard Index, Adamic-Adar (Resource Allocation), and Salton Index.
The results were stark: standard indices achieved a sensitivity of roughly 0.01. They simply couldn't handle the cross-domain jump. In contrast, the authors' approach, particularly when using QGrams for string matching, achieved a Precision of 0.908.
Table 1: Performance of various string similarity functions within the recursive framework.
Critical Insights & Future Outlook
The "Bridge-Assortativity" property discovered here—that i-bridges (users with me-edges) tend to congregate—is a powerful insight for digital forensics and marketing alike.
However, there is a clear Privacy-Utility Tradeoff. While this tool is a boon for data scientists wanting to create a "360-degree" user view, it represents a significant challenge for users attempting to maintain "security through obscurity" by keeping their professional (LinkedIn) and private (personal blog) personas separate.
Future work in this domain will likely need to incorporate Behavioral Semantics (analyzing what someone posts, not just who they follow) to identify users who intentionally choose completely different usernames and social circles across platforms.
Takeaway
For developers and researchers building multi-platform apps, this paper provides a robust mathematical framework to unify user identities without needing a "Login with Google" button, relying instead on the inherent structural properties of the social web.
