DualLink: Tackling the Heterogeneity Challenge in User Identity Linkage via Dual Domain Adaptation

DualLink: Dual Domain Adaptation for User Identity Linkage Across Social Networks

2021-01-01
Bei Xu, Yue Kou, Guangqi Wang, Derong Shen, Tiezheng Nie
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces DualLink, a novel User Identity Linkage (UIL) framework designed to connect accounts across different social networks using a Dual Domain Adaptation mechanism. It effectively combines adversarial learning for node embedding and back-propagation neural networks for optimized node matching, achieving state-of-the-art results on datasets like Twitter-Foursquare and Douban-Weibo.

TL;DR

Connecting user identities across disparate social platforms (e.g., matching a Twitter handle to a Foursquare account) is notoriously difficult due to "Domain Shift"—the difference in data distribution and network structure between platforms. DualLink addresses this by applying domain adaptation twice: first during node embedding via adversarial training to create "network-blind" features, and second during node matching using Back-Propagation Neural Networks (BPNN) to refine the alignment function.

Context: This work moves beyond single-step adaptation, establishing a more holistic pipeline for cross-network user alignment.

The Core Challenge: Inconsistent Distributions

In the world of social networks, no two "ecosystems" are the same. Twitter facilitates short-text broadcasting, while Foursquare focuses on location-based check-ins. When we attempt to represent users in these networks as vectors (embeddings), the resulting distributions are often shifted. Prior works typically focused on capturing local topology but ignored the fact that the mapping function itself might need to adapt to bridge the gap between Network A and Network B.

Methodology: The Dual-Stage Breakthrough

The architecture of DualLink is divided into two distinct adaptation stages designed to minimize the impact of network heterogeneity.

1. Stage One: Adversarial Node Embedding

DualLink uses a Deep Network Embedding module consisting of two parallel extractors:

  • FE1 (Node Attributes): Captures profile-centric data.
  • FE2 (Neighbor Attributes): Captures the topological context by aggregating the features of a user’s social circle.

To make these embeddings "network-invariant," the model introduces an Adversarial Game. A domain discriminator tries to guess which network a user belongs to, while the embedding module is trained to fool the discriminator. The result is a latent space where a user from Twitter and a user from Foursquare look like they belong to the same distribution.

Overall architecture of the DualLink model

2. Stage Two: Back-Propagation Node Matching

Even with aligned embeddings, the final matching requires a sophisticated function. DualLink employs a BPNN to learn the optimal cosine-distance mapping between anchor nodes (users known to be the same across platforms). By iteratively updating via Stochastic Gradient Descent (SGD), the model learns a robust matching function that compensates for any remaining domain residuals.

Experimental Performance

The researchers evaluated DualLink against heavyweights like DeepLink and PALE using the AMiner dataset (Twitter-Foursquare and Douban-Weibo).

Key Insights from Results:

  • Dimensionality Resilience: While some methods fail in higher dimensions, DualLink thrives. It showed significant performance gains as the vector dimensionality reached 100-128, indicating it captures more nuanced cross-network information.
  • Data Efficiency: Even with a small percentage of "anchor links" (training data), DualLink maintains a competitive F1-score, making it practical for real-world scenarios where labeled data is scarce.

Performance Comparison across different metrics

Ablation Study: Why it Works

The authors performed an ablation study (removing components one by one). They found that:

  1. Removing the Domain Discriminator (no-D) caused a sharp drop in Precision and Recall.
  2. Removing the Neighbor Extractor (FE2) significantly harmed the F1-score, highlighting that social topology is just as vital as user profiles.

Ablation Study Results

Critical Analysis & Conclusion

DualLink represents a shift from "feature engineering" to "domain alignment." By treating the UIL problem as a transfer learning task, the authors successfully mitigate the noise introduced by heterogeneous network environments.

Takeaway: The power of DualLink lies in its "Dual" nature. You cannot solve UIL by just having better features; you must also ensure those features speak the same "language" across domains and that your matching engine is tuned to that language.

Future Work: While effective, the current model relies on the existence of some anchor links. Future research could explore unsupervised dual domain adaptation to link users when no prior matching data is available.

Find Similar Papers

Try Our Examples

  • Find recent papers on User Identity Linkage (UIL) that utilize Graph Convolutional Networks (GCN) or Transformers for cross-domain alignment.
  • Which paper first introduced the concept of Adversarial Domain Adaptation for graph representation learning, and how does DualLink's architecture extend it?
  • Explore research studies that apply dual domain adaptation techniques to multi-modal entity alignment tasks, such as linking user identities across audio and video social platforms.
Contents
DualLink: Tackling the Heterogeneity Challenge in User Identity Linkage via Dual Domain Adaptation
1. TL;DR
2. The Core Challenge: Inconsistent Distributions
3. Methodology: The Dual-Stage Breakthrough
3.1. 1. Stage One: Adversarial Node Embedding
3.2. 2. Stage Two: Back-Propagation Node Matching
4. Experimental Performance
4.1. Key Insights from Results:
4.2. Ablation Study: Why it Works
5. Critical Analysis & Conclusion