CP-Link: Bridging Human Identities Across Social Networks via Continuous Spatio-Temporal Patterns

User Identity Linkage across Location-Based Social Networks with Spatio- Temporal Check-in Patterns

2020-12-01
Fengxiang Ding, Xiaoqiang Ma, Yang Yang, Chen Wang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces CP-Link, a novel User Identity Linkage (UIL) framework designed to match accounts across different Location-Based Social Networks (LBSNs). By leveraging continuous spatio-temporal check-in patterns and a user-associated frequent pattern model, it achieves a significant performance leap, outperforming state-of-the-art methods by over 20% in AUC.

TL;DR

Connecting the dots between a Foursquare check-in and a Twitter post by the "same person" is a notorious challenge due to data sparsity and "the grid problem." CP-Link solves this by abandoning rigid spatial bins in favor of continuous stay regions and a time-series similarity engine (IDTW), resulting in a 20%-40% boost in AUC over established SOTA methods.

The "Discretization" Trap: Why Grids Fail

In User Identity Linkage (UIL), the goal is to prove that User A on Platform X is the same as User B on Platform Y. Most researchers try to find "encounters"—occasions where both accounts "check in" at the same time and place.

However, the common practice of dividing the world into grids or bins creates two fatal flaws:

  1. Information Loss: If a user checks in right on the edge of two grid cells, the system might fail to link them.
  2. Rigidity: Human behavior isn't "binned." Our movements are continuous paths. Discretizing these paths destroys the subtle periodic patterns that make our mobility signatures unique.

Methodology: The CP-Link Architecture

The authors propose a dual-engine system to overcome these hurdles: data enrichment and continuous matching.

1. The LFP-Tree: Solving Sparsity with "Friends"

If a user only has three check-ins, you can't build a reliable pattern. The Location Frequent Pattern (LFP) model identifies "associated users"—others with similar mobility—and uses their frequent locations to "fill in the blanks" for the target user. This creates a denser, more actionable dataset without inventing synthetic noise.

System Architecture

2. Density-Peaks (DP) Stay Regions

Instead of squares on a map, CP-Link finds Density Peaks. It identifies locations where a user spends significant time and groups nearby points into irregularly shaped regions. This respects the physical reality of "staying" (e.g., a home or office) better than a predefined grid.

3. IDTW: Time-Series Similarity

Once stay regions are identified, the "check-in" sequence within each region is treated as a time series. The authors use an Improved Dynamic Time Warping (IDTW) algorithm. Unlike a simple distance check, IDTW can:

  • Align sequences of different lengths.
  • Focus on the Longest Common Subsequence (LCS) to find periodic behaviors (e.g., visiting the gym every Tuesday).
  • Assign weights based on the importance of specific stay regions.

Experimental Results: Stability is King

The most striking result of CP-Link isn't just the high AUC; it's the robustness.

Performance Comparison

As shown in the charts, traditional methods (POIS, SIMP) see their F1 scores fluctuate wildly as you change the "grid size." CP-Link stays flat and superior. Because it doesn't rely on grids, the spatial resolution of the data doesn't "break" the algorithm.

MetricCP-Link (Proposed)GKR-KDE (SOTA)Improvement
AUC~0.9~0.65+40%
PrecisionHigher/StableVariableSignificant

Critical Insight & Conclusion

CP-Link proves that User Identity Linkage is a "Sequence" problem, not a "Coordinate" problem. By stoping the obsession with "where" a user is exactly at 12:00 PM and starting to look at the "rhythm" of their movements across stay regions, we can bridge accounts even when data is sparse.

Future Outlook: While CP-Link is powerful, its reliance on finding "associated users" might be computationally expensive as platforms scale to billions of users. Integrating this continuous logic into a Graph Neural Network (GNN) could be the next frontier to further automate feature extraction and improve efficiency.

Find Similar Papers

Try Our Examples

  • Find recent papers published after 2020 that utilize Graph Neural Networks (GNNs) or Hypergraph Embeddings specifically for User Identity Linkage in check-in data.
  • Which study first introduced the concept of 'Stay Region' in trajectory mining, and how does the Density-Peaks approach in this paper structurally improve upon that original definition?
  • Investigate how CP-Link's continuous spatio-temporal modeling could be adapted for cross-domain privacy-preserving tasks, such as differential privacy in trajectory synthesis.
Contents
CP-Link: Bridging Human Identities Across Social Networks via Continuous Spatio-Temporal Patterns
1. TL;DR
2. The "Discretization" Trap: Why Grids Fail
3. Methodology: The CP-Link Architecture
3.1. 1. The LFP-Tree: Solving Sparsity with "Friends"
3.2. 2. Density-Peaks (DP) Stay Regions
3.3. 3. IDTW: Time-Series Similarity
4. Experimental Results: Stability is King
5. Critical Insight & Conclusion