Identifying Users Across Social Media: The Hidden Signal in Display Names
User Identification Based on Display Names Across Online Social Networks
This paper introduces a machine learning-based framework for identifying the same user across multiple Online Social Networks (OSNs) using only "Display Names." By leveraging a set of 14 string-based features and Logistic Regression with cross-validation (LRCV), the authors achieve SOTA F1-scores of up to 96.24% on real-world datasets from Facebook, Twitter, and Foursquare.
TL;DR
Digital identity is fragmented across social platforms. While most research uses private data or complex social graphs to link accounts, this paper proves that publicly available "Display Names" are enough. By using a machine learning model built on string redundancy and naming patterns, the authors achieved over 96% accuracy in linking accounts across Facebook, Twitter, and Foursquare.
Background: The Cost of Data Acquisition
In the academic landscape of user identification, we usually see two extremes:
- Multi-dimensional methods: These are accurate but intrusive, requiring access to friend lists, location history, and private profiles.
- Username-based methods: These fail because many sites (like Foursquare or QQ) assign arbitrary numeric strings as usernames.
This paper identifies a "sweet spot": the Display Name. It is public, low-cost to acquire, and—critically—constrained by human psychology. Humans tend to reuse or slightly vary their names because of memory limits, a phenomenon known as Information Redundancy.
Methodology: Capturing the Human "Naming Logic"
The authors propose a supervised learning framework. The core innovation lies in the 14 features designed to quantify how "similar" two names are, beyond simple spelling.
1. The Similarity Matrix (Average of Best Match)
Unlike simple string matching, the authors segment names into word arrays.
- Example: "David J. Whelan" vs "Dave Whelan".
- They calculate a similarity matrix between each word and use a step-wise reduction to find the best pairings, capturing shifts like "David" "Dave" or the omission of initials.
2. Information Redundancy Features
- SimLCS: Measures the longest common substring relative to total length.
- Jensen-Shannon Distance: Analyzes the distribution of characters. Even if the order is flipped (e.g., "Gateman" and "Nametag"), the character distribution remains nearly identical.
- Normalized Edit Distance: Measures the minimum operations (insert, delete, substitute) to turn one name into another.

Experimental Results & Baseline Comparison
The authors crawled nearly 600,000 profiles. They compared their method against two major baselines: MOBIUS (a behavioral modeling approach) and Liu's approach (a username-based feature set).
| Dataset | Method | Accuracy | F1-Score |
|---|---|---|---|
| FB-FS | Ours (LRCV) | 96.24% | 96.24% |
| MOBIUS | 95.36% | 95.36% | |
| Liu's | 90.04% | 89.32% |
The "Identical Name" Challenge
Critics might argue that identification is easy because people use the exact same name. The authors addressed this by conducting an ablation study on a non-identical dataset. Even when identical names were removed, their model remained superior, dropped only slightly to ~85-90% F1-score, whereas simpler models collapsed.

Deep Insights: Privacy and Robustness
The feature importance analysis (using Mutual Information and Random Forests) revealed that the naming patterns (Order of names, common substrings) are the most predictive.
Takeaway for Privacy: If you want to remain unlinked, simply changing a "middle name" or "nicknaming" yourself is insufficient. The underlying character distribution and partial substring overlaps are strong signals.
Conclusion & Limitations
This work demonstrates that user identification doesn't require "big data"—it just requires "smart features" applied to public data. However, as the authors note, newer users are becoming more privacy-conscious, leading to greater name divergence. Future research will likely need to incorporate "Latent Space" modeling to handle cases where names provide zero overlap.
Why this matters:
For business intelligence and fraud detection, this offers a high-speed, low-cost way to link identities without violating the Terms of Service of platforms that restrict private data harvesting.
