Bridging the Social Gap: A Hybrid Framework for Cross-Platform Seed User Identification
A framework for seed user identification across multiple online social networks
This paper introduces a hybrid personal information-based framework for Seed User Identification (SUI) across multiple social platforms, specifically Twitter and Instagram. By combining "cross-link posts" (content) with "following network" relations (structure), the method identifies core seeds to facilitate large-scale user mapping.
TL;DR
As users migrate between platforms like Twitter and Instagram, they leave digital breadcrumbs. This paper presents a framework that uses cross-link URLs and following network structures to identify "Seed Users"—the bridge nodes that allow us to link identities across the social web. By matching network neighbors using Levenshtein distance, the authors found that a single link can reveal dozens of connected accounts.
Problem & Motivation: The Identity Fragmentation
The modern social media landscape is fragmented. A user might be a professional on LinkedIn, a visual storyteller on Instagram, and a news commentator on Twitter. Identifying that these accounts belong to the same person is a "hard nut to crack" due to:
- Platform Heterogeneity: Different data structures and privacy rules.
- Identity Impersonation: Similar usernames belonging to different people (e.g., the Sonu Sood vs. Sonu Nigam incident).
- Data Sparsity: Short-form content makes traditional linguistic stylometry difficult.
The authors' key insight is Social Mirroring: users typically replicate their real-world social circles across platforms. If User A follows a specific set of people on Twitter, they are highly likely to follow those same individuals on Instagram.
Methodology: The Two-Step Bridge
The framework operates in two distinct phases:
Phase 1: Seed Identification via Cross-Links
The process begins by searching for "Self-Mentioned URLs." These are posts where a Twitter user explicitly shares their Instagram handle (e.g., "Follow me on Instagram: @username"). This provides a 100% accurate ground-truth link for an initial set of users.
Phase 2: Network Alignment
Once an initial link is established, the algorithm crawls the following lists of that user on both SN1 (Twitter) and SN2 (Instagram).
- Levenshtein Distance: The system compares the usernames in both "Following" sets.
- String Matching: Usernames with zero edit distance (exact matches) are flagged as potential seed users.
- Validation: These matches are verified as "Final Seeds" to be used for further network propagation.
Figure 1: Conceptual representation of identifying common users across network boundaries.
Experiments & Results
The study analyzed 2,193 profiles on Twitter and 952 on Instagram derived from cross-linked posts.
Key Findings:
- Efficiency: On average, one cross-link post yields 52 seed users.
- Consistency: Manual verification confirmed that users with identical usernames across platforms almost always shared other attributes like profile pictures and descriptions.
- Scalability: The method works even when profiles are partially private, as long as the "following" list is accessible via API.
Table 1: The distribution of following profiles and identified identical seeds for sample users.
Critical Analysis & Conclusion
Takeaway
The value of this work lies in its Preprocessing Efficiency. Rather than attempting to match millions of users blindly, this framework targets "hubs" (seeds) that naturally link the two networks. It moves beyond fragile profile metadata into the more stable territory of social topology.
Limitations
- Attribute Sensitivity: If a user adopts drastically different naming conventions across platforms (e.g., "TechExplorer" on Twitter vs. "JohnDoe88" on Instagram), the Levenshtein distance of zero will fail.
- Privacy Trends: As platforms increasingly restrict API access to "following" lists (shadowing Facebook's post-Cambridge Analytica shift), the structural data required for this method may become harder to extract.
Future Outlook
The next frontier for this research involves using Jaccard Similarity and State-Space Models to handle near-match usernames and evolving social graphs. This framework provides the foundational "hooks" necessary for any robust multi-platform analytics engine.
