The Digital Echo: Why Your Display Names Are More Traceable Than You Think
Understanding the User Display Names across Social Networks
This paper presents a comprehensive measurement study on user display names across Foursquare, Facebook, and Twitter, introducing a distributed crawling framework to analyze "redundant information." The study discovers that over 45% of users maintain identical display names across platforms, achieving a state-of-the-art behavioral profile for cross-platform user identification.
TL;DR
Researchers from Northwestern Polytechnical University have quantified the "redundancy" in how we name ourselves online. By analyzing 1.3 million accounts across Facebook, Twitter, and Foursquare, they found that nearly half of us use identical names everywhere, and even those who don't leave behind a "letter distribution" fingerprint that is statistically distinct from other users.
The Motivation: The Memory Trap
Why do we use the same names? The study pivots on a psychological Inductive Bias: Human memory limitations. We tend to use consistent behaviors to make our digital presence easier to remember or to build a unified online reputation.
While previous research heavily focused on usernames (the unique, often alphanumeric handles), this paper argues that display names (the "friendly" names) are actually more valuable. In many networks like Foursquare or QQ, usernames are just random strings of numbers, making display names the primary conduit for human-readable identity.
Methodology: Mining the Cross-Site Link
The authors exploited the "Cross-site linking" function of Foursquare—where users voluntarily link their Facebook and Twitter profiles. This provided a rare "Golden Dataset" where identities are already confirmed, allowing for a rigorous comparison between "Positive" pairs (the same person) and "Negative" pairs (different people).

The Three Dimensions of Similarity
To quantify just how similar these names are, the team applied three specific metrics:
- Character Similarity: Using LCS (Longest Common Substring) and Edit Distance. They found that if the edit distance between two names is less than half the total length, there is a very high probability they belong to the same person.
- Visual Structure (Best Match): Segmenting names into <First, Middle, Last> and checking if parts are swapped or omitted.
- Statistical Fingerprinting: Using Jensen-Shannon (JS) Divergence to compare the probability distribution of the 26 English letters. Even if you change "John Doe" to "Doe John," your JS similarity remains nearly 1.0.
Key Insights from the Data
The study revealed several SOTA observations regarding user behavior:
- The 45% Rule: Between 45% and 63% of individuals use the exact same display name across platforms.
- Real-Life Mirroring: The letter distribution in OSN display names almost perfectly mirrors the distribution of common names in real-life census data (peaks at 'a', 'e', 'n', 'i').
- Platform Specificity: Names on Facebook and Foursquare are more similar to each other than to Twitter, likely because the former two encourage "real-name" usage while Twitter leans toward "handles."

Persistence over Time
One of the most significant contributions of this work is the Evolution Analysis. By dividing the data into nine chronological chunks based on registration ID, the authors proved that these naming behaviors are time-independent. Whether you joined a network in 2010 or 2016, the redundancy of your display names remains constant.

Critical Analysis & Takeaways
This paper provides the mathematical "glue" for future cross-site user identification systems.
Takeaway 1: Privacy is an illusion of effort. Most users who try to change their names for privacy only change parts of them (e.g., omitting a last name), which does little to stop modern string-matching algorithms.
Takeaway 2: Letter distribution is an underrated feature. The fact that JS similarity for positive pairs is consistently above 0.8 while negative pairs stay below 0.7 suggests that "letter frequency" can serve as a lightweight pre-filter for large-scale identity matching.
Limitations: The study focuses primarily on the 26-letter English alphabet. As social networks grow in non-Western regions, the metrics of LCS and Edit Distance may need to be adapted for multibyte character sets or phonetic transcriptions of non-Latin names.
