LHNE: Synthesizing Structure and Content for Seamless User Identity Linkage
User identity linkage across social networks via linked heterogeneous network embedding
The paper introduces LHNE (Linked Heterogeneous Network Embedding), a model for User Identity Linkage (UIL) across social networks. It bridges heterogeneous networks by jointly embedding structural (friendship links) and content (user-generated topics) information into a unified low-dimensional latent space, achieving state-of-the-art performance on Twitter-Flickr and DBLP datasets.
TL;DR
Connecting personal identities across different social platforms (e.g., linking a Twitter handle to a Flickr account) is a cornerstone of modern cross-network analytics. LHNE (Linked Heterogeneous Network Embedding) provides a breakthrough by refusing to treat social links and post content as separate entities. Instead, it embeds them into a unified "Interest-Friendship" manifold, resulting in a 47%+ performance boost over previous state-of-the-art models.
The "Independence" Fallacy in Prior Work
Most researchers previously assumed that "who you follow" (structure) and "what you say" (content) are independent variables. They would build one model for the social graph and another for text, then stitch them together like a Frankenstein’s monster.
However, the authors of LHNE identify a crucial Correlation Insight: A user following a celebrity is statistically likely to post about that celebrity. By failing to model this correlation in a shared space, previous methods lost critical information, especially for "isolated" users who lack social edges but are active content creators.
Methodology: The Unified Latent Space
LHNE addresses the heterogeneity problem through a multi-stage pipeline:
1. Topic-Based Denoising
Social media text is noisy (slang, ads, typos). LHNE uses Latent Dirichlet Allocation (LDA) to extract "Topics of Interest." This transforms raw text into a stable probability distribution over themes, effectively acting as a filter for irrelevant noise.
2. The Four Pillars of Linking
The model constructs a linked heterogeneous network consisting of four distinct sub-networks:
- User-User Intra-network: Friendships within a platform.
- User-Topic Intra-network: Interests within a platform.
- User-User Inter-network: Transferring identity via known "anchor" users.
- User-Topic Inter-network: Aligning topics across platforms (e.g., "Photography" on Flickr vs. "Camera" on Twitter).

3. Joint Embedding Learning
The core of LHNE is its objective function. It doesn't just minimize the distance between friends; it minimizes the KL-divergence across all four sub-networks simultaneously. Using Negative Sampling, it ensures that "User A" from Twitter and "User A" from Flickr converge to the same point in a low-dimensional vector space.
Experimental Results: Dominance in Sparsity
The researchers tested LHNE against strong baselines like IONES and KNN on a Twitter-Flickr dataset.
Key Findings:
- Strength in Isolation: In Flickr, roughly 40% of users had no friendship links. Standard structure-only models failed here. LHNE used their content "topics" as a proxy context to successfully identify them.
- Efficiency: LHNE converges faster and requires fewer dimensions () to reach stability compared to its predecessors.
Figure: LHNE showing superior Recall and Precision as the similarity threshold changes.
Critical Analysis: Why This Matters
The true value of LHNE lies in its robustness to data difficulty. In a world of increasing privacy protections and API limitations, we often cannot get a full social graph. LHNE proves that if you have even a sliver of content data, you can reconstruct the missing social context via the topic manifold.
Limitations: While powerful, the model relies on a "Seed Set" of known anchor users. In a purely cold-start environment where NO links are known, the inter-network transfer might struggle. Future work could look into zero-shot alignment using cross-lingual or cross-modal embeddings.
Conclusion
LHNE represents a shift from "Multi-view learning" (looking at different features separately) to "Unified Manifold Alignment." By recognizing that our interests and our friends are two sides of the same coin, it provides the most accurate bridge yet between our fragmented digital selves.
