Bridging the Digital Divide: Recursive Identity Discovery Across Social Networks

Discovering missing me edges across social networks q

2015-05-18
Francesco Buccafurri, Gianluca Lax, Antonino Nocera, Domenico Ursino
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a specialized method for discovering missing "me edges"—links between accounts belonging to the same individual—across disparate social networks. It proposes a recursive common-neighbor approach combined with string similarity that achieves high precision in real-life Social Internetworking Scenarios (SIS).

TL;DR

As users fragment their digital lives across platforms like Twitter, LinkedIn, and Instagram, the "me edge"—a formal link between two accounts owned by the same person—is frequently missing. This paper proposes a recursive, node-similarity algorithm that overcomes the limitations of traditional link prediction. By combining username string analysis with a specialized "bridge-assortativity" heuristic, the authors can identify missing identity links with over 90% precision, effectively unifying fragmented social graphs.

The Identity Fragmentation Problem

In the era of Social Internetworking Scenarios (SIS), we are no longer defined by a single social graph but by a "constellation" of networks. While standards like FOAF (Friend-Of-A-Friend) and XFN allow users to explicitly declare "this is me" on another site, most users neglect to do so.

This creates a massive "missing link" problem. Traditional Social Network Analysis (SNA) tools are designed for single-network silos. When you try to find a user across two different networks, standard Common-Neighbor algorithms fail because the neighbors themselves are "anonymous" relative to the other network. It's a recursive cold-start problem: you can't find the user's identity without knowing their friends' identities, but you can't know the friends' identities without a unified graph.

Methodology: Recursive Similarity & Public Figure Filtering

The authors' solution is an elegant recursive operator . Instead of a static check, the algorithm propagates similarity through the network layers.

1. The Similarity Operator

The similarity of two accounts is defined by:

  • String Similarity: Using metrics like QGrams or Levenshtein to compare usernames (e.g., johnsmith vs dj_smith).
  • Weighted Neighbor Overlap: A recursive step where the similarity of friends' accounts contributes to the root pair's score.

2. The "Public Figure" Problem

A major pitfall in graph analysis is the "Obama Effect." If two people both follow a global celebrity, it doesn't mean they are the same person; it just means they are both on the internet. The paper introduces a scaling coefficient that diminishes the weight of accounts with extremely high indegree, effectively filtering out the noise of mass-followings.

Model Logic and Flow Figure 1: The Bridge-Assortativity logic, finding candidates by traversing existing me-edges of neighbors.

Experiments: Why Traditional SOTA Fails

The authors pitted their method against established indices like the Jaccard Index, Adamic-Adar (Resource Allocation), and Salton Index.

The results were stark: standard indices achieved a sensitivity of roughly 0.01. They simply couldn't handle the cross-domain jump. In contrast, the authors' approach, particularly when using QGrams for string matching, achieved a Precision of 0.908.

Experimental Results Table Table 1: Performance of various string similarity functions within the recursive framework.

Critical Insights & Future Outlook

The "Bridge-Assortativity" property discovered here—that i-bridges (users with me-edges) tend to congregate—is a powerful insight for digital forensics and marketing alike.

However, there is a clear Privacy-Utility Tradeoff. While this tool is a boon for data scientists wanting to create a "360-degree" user view, it represents a significant challenge for users attempting to maintain "security through obscurity" by keeping their professional (LinkedIn) and private (personal blog) personas separate.

Future work in this domain will likely need to incorporate Behavioral Semantics (analyzing what someone posts, not just who they follow) to identify users who intentionally choose completely different usernames and social circles across platforms.

Takeaway

For developers and researchers building multi-platform apps, this paper provides a robust mathematical framework to unify user identities without needing a "Login with Google" button, relying instead on the inherent structural properties of the social web.

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend cross-platform user identity resolution using graph neural networks (GNNs) or embedding-based manifold alignment.
  • Which study first introduced the concept of "me edges" in the context of the XHTML Friends Network (XFN), and how has the formal definition evolved in current Social Internetworking Scenarios?
  • Explore research that applies bridge-assortativity or membership overlap analysis to detect Sybil attacks or coordinated inauthentic behavior across multiple social media platforms.
Contents
Bridging the Digital Divide: Recursive Identity Discovery Across Social Networks
1. TL;DR
2. The Identity Fragmentation Problem
3. Methodology: Recursive Similarity & Public Figure Filtering
3.1. 1. The Similarity Operator
3.2. 2. The "Public Figure" Problem
4. Experiments: Why Traditional SOTA Fails
5. Critical Insights & Future Outlook
5.1. Takeaway