DBUL: Reconsidering User Identity Linkage through the Lens of Density-Based Clustering

DBUL: A User Identity Linkage Method across Social Networks Based on Spatiotemporal Data

2021-11-01
Hui Xue, Bo Sun, Chengxiang Si, Wei Zhang, Jing Fang
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces DBUL, a novel User Identity Linkage (UIL) method using DBSCAN clustering to match users across social networks via spatiotemporal data (FS-TW and IG-TW). By representing scattered check-in points as distinct cluster centers, the method achieves SOTA precision and significantly improves computational efficiency compared to grid-based or trajectory-based baselines.

TL;DR

Connecting the dots between a Foursquare check-in and a Twitter status update isn't just a matter of drawing lines. DBUL (DBSCAN-based User Linkage) revolutionizes this by discarding the rigid "trajectory" and "grid" models. Instead, it uses density-based clustering to extract the "essence" of a user's location habits, delivering SOTA precision and reducing computational overhead by up to 46.5%.

The "Sparsity" Trap in Spatiotemporal Data

User Identity Linkage (UIL) is the backbone of cross-domain recommendation and personalized advertising. However, social network spatiotemporal data is notoriously messy:

  • Sparsity: Unlike a continuous vehicle GPS trace, social check-ins are snapshots separated by hours or even days.
  • Grid Distortions: Previous SOTA methods like GKR-KDE chop the world into grids. If a user check-ins on the very edge of two adjacent grids, the system fails to recognize them as the same location—the "edge exception" problem.
  • Noise: Random one-time check-ins often skew trajectory-matching algorithms.

The authors of DBUL propose a simple yet powerful physical intuition: Humans are creatures of habit. Regardless of which app they use, they gravitate toward a few "cluster centers" (home, office, favorite cafe).

Methodology: From Points to Cluster Centers

DBUL replaces raw data points with a set of cluster centers .

1. The DBSCAN Advantage

Unlike K-Means, DBSCAN doesn't require knowing the number of clusters in advance and can handle irregular shapes. This is perfect for capturing the unique spatial "footprint" of a human user.

2. The Similarity Architecture

The method employs a modified Ochiai Coefficient to calculate similarity. Because two GPS coordinates rarely match perfectly, it introduces a coincidence_radius. If two cluster centers from different platforms are within this radius, they are counted as an intersection.

Model Architecture

3. Smart Noise Handling

The paper introduces "Reserve Conditionally" logic. If a user's data is overwhelmingly noisy, the noise points are retained as individual clusters to prevent total information loss. If clear clusters exist, the noise is discarded to maintain efficiency.

Experimental Battleground: FS-TW vs. IG-TW

The researchers tested DBUL against industrial-strength baselines including BIN, DG, and GKR-KDE.

Performance Highlights:

  • Precision Supremacy: On the Foursquare-Twitter (FS-TW) dataset, DBUL hit the highest precision among all methods.
  • Density Robustness: On the larger Instagram-Twitter (IG-TW) dataset, which is denser and more complex, DBUL swept all metrics (Precision, Recall, F1).

Performance Results

Efficiency Gains:

By reducing thousands of records into a handful of cluster centers, the computational volume drops drastically. DBUL clocked an average running time significantly lower than BIN and GS, making it a viable candidate for real-time large-scale deployments.

Critical Insight & Conclusion

The true value of DBUL lies in its Inductive Bias. By assuming that user behavior is center-driven rather than pathway-driven, it finds a "signal" in the "noise" of sparse social data.

Takeaway: For practitioners dealing with sparse behavioral data, the lesson is clear: don't try to reconstruct the path; find the hubs. While DBUL currently focuses on spatial data, its clustering logic could easily be extended to temporal or even semantic behavioral clusters in future iterations.

Find Similar Papers

Try Our Examples

  • Search for recent papers published after 2020 that address User Identity Linkage in sparse spatiotemporal datasets using Graph Neural Networks or Contrastive Learning.
  • Identify the seminal paper that established the Ochiai coefficient in ecological studies and trace its adaptation into modern data mining for similarity measurement.
  • Explore how density-based clustering models like DBUL can be integrated into privacy-preserving frameworks for federated user identity matching.
Contents
DBUL: Reconsidering User Identity Linkage through the Lens of Density-Based Clustering
1. TL;DR
2. The "Sparsity" Trap in Spatiotemporal Data
3. Methodology: From Points to Cluster Centers
3.1. 1. The DBSCAN Advantage
3.2. 2. The Similarity Architecture
3.3. 3. Smart Noise Handling
4. Experimental Battleground: FS-TW vs. IG-TW
4.1. Performance Highlights:
4.2. Efficiency Gains:
5. Critical Insight & Conclusion