Linking Identities: Social Summarization for Cross-Platform Person Identification
Person Identification between Different Online Social Networks
This paper introduces a Social Summary method for Person Identification across heterogeneous Online Social Networks (OSNs). By employing a two-phase clustering approach to model an individual's social circle, the method achieves superior cross-platform matching performance, significantly outperforming traditional profile and username comparison baselines on real-world Facebook and Wretch datasets.
TL;DR
Identifying the same individual across different Online Social Networks (OSNs) is a daunting task due to privacy settings, missing profile data, and varying naming conventions. This paper presents a novel Social Summary framework that moves beyond simple text matching. By clustering a user's friends into cohesive social groups and analyzing the aggregate attributes of these groups, the authors achieve highly accurate person identification even when primary profile data is sparse or inconsistent.
The "Digital Ghost" Problem: Why Simple Matching Fails
In an ideal world, cross-network identification is trivial—look for the same email or username. However, in reality:
- Missing Values: Users are lazy or privacy-conscious; profiles are often incomplete.
- The Cross-Cultural Gap: On Facebook, a user might use "John Doe," but on a local platform like Wretch, they use "張小明." Standard string metrics like Jaro distance fail here.
- Computational Explosion: Comparing every friend of User A with every friend of User B is an nightmare that doesn't scale.
The authors' core insight is based on Homophily: Birds of a feather flock together. Even if a user hides their details, their social clusters (e.g., high school friends, coworkers) remain consistent across platforms.
Methodology: The Two-Phase Social Summary
The proposed method replaces one-to-one friend comparison with a distribution-based approach.
1. Friend Clustering (Filtering the Noise)
Instead of treating the friend list as a flat list, the algorithm identifies "social seeds" by iteratively removing critical nodes (bridges between groups) using Betweenness Centrality. This reveals the dense cores of different social circles (e.g., family vs. hobbyists).
2. Social Summary Generation
Once clusters are formed, the system generates a summary for each using TF-ICF (Term Frequency - Inverse Cluster Frequency). This highlights attributes that are unique to a specific group, creating a "signature" for that social circle.
(Note: This figure would illustrate the transition from a raw friendship graph to refined seeds and finally to summarized clusters.)
3. Identity Resolution
Similarity is no longer just . It becomes a combination of:
- Direct Profile Similarity: Traditional attribute matching.
- Group Similarity: A max-match between the set of social summaries in Network A and Network B.
Experimental Battleground: Facebook vs. Wretch
The system was tested on a dataset of 100 ground-truth links hidden among 41,035 candidates.
Key Performance Metrics:
- Baseline (Username/Profile): Performed poorly due to the "cross-language" naming issue.
- Proposed Method: Achieved an MRR of 0.40, significantly higher than the 0.18 for name matching.
- Attribute Effectiveness: "Education" proved to be the strongest summary attribute, as users are most consistent about their schools across platforms.
(Note: This chart shows how the Social Summary method successfully shifts the majority of ground truths into the Top-100 tier compared to the erratic performance of basic methods.)
Critical Insight & Practical Value
The defining advantage of this work is its computational efficiency. By comparing a small number of clusters (usually < 10) rather than hundreds of individual friends, the system avoids the "computational explosion" of previous advanced methods.
However, a notable limitation is the reliance on friend lists being public. As platforms like Meta (Facebook) restrict API access to friend lists for privacy reasons, the "Social Summary" would need to adapt to use other signals like public interactions (likes, tags) or content-based "interest homophily."
Conclusion
This paper proves that our "social footprint"—the groups we belong to—is often more identifiable than our "profile face." For researchers in social graph analysis and cybersecurity, this work provides a robust blueprint for linking fragmented digital identities in an increasingly siloed internet.
