MLM: Bridging the Gap Between User Privacy and Cross-Platform Social Linkage

Differential Privacy-Preserving User Linkage across Online Social Networks

2021-06-25
Xin Yao, Rui Zhang, Yanchao Zhang
Summary
Problem
Method
Results
Takeaways
Abstract

The paper proposes a novel Differential Privacy (DP) framework for user linkage across Online Social Networks (OSNs) using a "Social Data Collector" (SDC) model. It introduces the Multivariate Laplace Mechanism (MLM) and two new privacy definitions—ε-attribute and ε-profile indistinguishability—to achieve high linkage accuracy while protecting sensitive user identities and attributes across platforms like Twitter, MySpace, and Last.fm.

TL;DR

Connecting a single user across Twitter, LinkedIn, and Yelp is a goldmine for recommendation systems but a nightmare for privacy. Standard Local Differential Privacy (LDP) often renders the data useless for such tasks. This paper introduces the Multivariate Laplace Mechanism (MLM), which uses correlated noise and distance-based indistinguishability to keep user profiles linkable for AI models while remaining mathematically private against identity leakage.

The "Curse" of Uniform Privacy

The central conflict in privacy-preserving data sharing is the Utility-Privacy Tradeoff. In User Linkage, a Social Data Collector (SDC) wants to know if "User A" on Twitter is "User B" on MySpace.

Standard Local Differential Privacy (LDP) is a "hard" constraint: it requires that the output for any two inputs be almost identical. If two users have vastly different attributes (e.g., one lives in New York, the other in London), LDP forces the system to add enough noise to make them look the same. In high-dimensional social profiles, this noise accumulates (the "curse of dimensionality"), making it impossible for a classifier to spot the subtle similarities that identify a single person.

The Insight: Distance-Based Indistinguishability

Instead of making everyone look the same, the authors argue that we only need to make similar people look the same. They define:

  • ε-Attribute Indistinguishability: If two users have similar ages, their perturbed ages should be indistinguishable.
  • ε-Profile Indistinguishability: If two users have similar overall profiles (Manhattan distance), their perturbed profiles should be indistinguishable.

By tying privacy strength to the distance between records, they allow the data to retain its "shape," allowing machine learning models to function effectively.

Methodology: The Multivariate Laplace Mechanism (MLM)

The core technical contribution is the Multivariate Laplace Mechanism (MLM). Unlike standard Laplace mechanisms that add noise to each attribute (age, location, etc.) independently, MLM adds positively correlated noise.

1. The Architecture

The workflow involves three distinct phases: Perturbation, Training, and Prediction.

Overall Framework Figure 1: The SDC collects perturbed data from OSNs, trains a classifier on a "consent" subset, and predicts links for anonymous users.

2. Why Correlated Noise?

If you add independent noise to 10 different attributes, the record "drifts" in 10 different directions, rapidly increasing the Manhattan distance from its original state. By adding correlated noise (controlled by a correlation coefficient ), the attributes drift together, preserving the internal logic of the profile used by classifiers like Logistic Regression or XGBoost.

Experimental Proof: Accuracy vs. Privacy

The authors tested their framework on Twitter, MySpace, and Last.fm datasets. The results were striking when compared to standard LDP mechanisms like Duchi or Piecewise Mechanism (PM).

Prediction Accuracy Figure 2: Precision on Twitter-MySpace. MLM (top lines) maintains significantly higher accuracy at low privacy budgets (small ε) compared to traditional LDP.

Key Findings:

  • High Precision at Low ε: At (strong privacy), MLM achieved ~63% precision, while others failed to beat 50%.
  • Sensitivity to ρ: As the correlation increased from 0.1 to 0.9, the prediction precision scaled up, proving that the correlation of noise is the "secret sauce" for data utility in linkage tasks.
  • Defense Capability: Despite the higher utility, the mechanism successfully reduced the "Inference Rate" (the ability of an attacker to identify a user) compared to raw, unperturbed data.

Critical Perspective

While MLM is a major step forward, it relies on a semi-trusted SDC and assumes that a subset of users will "volunteer" their raw data for training. In a real-world scenario, the incentive for users to waive their privacy for "online credits" might be low, potentially biasing the training set.

Conclusion

This paper demonstrates that privacy doesn't have to mean "useless data." By moving away from the rigid boundaries of standard LDP and embracing the natural correlations within multidimensional data via the Multivariate Laplace Mechanism, we can protect users' identities without blinding the algorithms that make social networks useful.

Find Similar Papers

Try Our Examples

  • Find recent papers that extend distance-based differential privacy (like geo-indistinguishability) to graph-structured social network data or node linkage tasks.
  • Which study first introduced the concept of adding correlated noise in Local Differential Privacy, and how does the Multivariate Laplace Mechanism compare to modern Matrix Mechanism approaches?
  • Search for research applying Multivariate Laplace Mechanisms or similar correlated noise techniques to privacy-preserving federated learning or multi-modal data fusion.
Contents
MLM: Bridging the Gap Between User Privacy and Cross-Platform Social Linkage
1. TL;DR
2. The "Curse" of Uniform Privacy
3. The Insight: Distance-Based Indistinguishability
4. Methodology: The Multivariate Laplace Mechanism (MLM)
4.1. 1. The Architecture
4.2. 2. Why Correlated Noise?
5. Experimental Proof: Accuracy vs. Privacy
5.1. Key Findings:
6. Critical Perspective
7. Conclusion