Linking Shadows: High-Precision User Identification Across LinkedIn and Twitter
Similarity-Based User Identification Across Social Networks
This paper presents a similarity-based framework for identifying users across different social networks (specifically LinkedIn and Twitter) using a trainable combination of metrics. By aligning professional data, location, and social interactions, the authors achieve high-accuracy cross-platform linking (Recall up to 94.27%) to support information verification.
TL;DR
As our digital footprints scatter across platforms, the ability to "bridge" identities becomes crucial for verifying information. This paper introduces a supervised learning framework that combines professional achievements, geospatial data, and string similarity to identify the same individual across LinkedIn and Twitter with over 94% accuracy, effectively solving the name disambiguation problem in a professional context.
Problem & Motivation: The "John Smith" Paradox
In the context of the REVEAL project—aimed at verifying the trustworthiness of social media contributors—the biggest hurdle is Presence. If a journalist sees a claim on Twitter, how can they verify the author's credentials on a professional platform like LinkedIn?
The problem isn't just finding a name; it's disambiguation. A search for a common name returns dozens of profiles. Prior work often struggled with:
- Data Sparsity: Users leave many profile fields (location, bio) blank.
- Class Imbalance: In a set of 25 search results, only one (or zero) is a true match, making it easy for classifiers to "favor" the mismatch class.
Methodology: The Architecture of Similarity
The authors transform the identification task into a binary classification problem (Match vs. Mismatch). They generate a Similarity Vector based on five specialized dimensions.

1. The Five Pillar Metrics
- Name (Jaro-Winkler): Handles slight variations in name spelling.
- Description (Token Ratio): Compares short biographies by calculating the ratio of common keywords.
- Location (Geospatial Semantic): Uses the GeoNames ontology. Instead of raw text matching, it calculates the ratio of "bounding box" areas. If someone says "Manhattan" on Twitter and "New York" on LinkedIn, the system recognizes the spatial subsumption.
- Affiliation-Education (Smith-Waterman): Maps LinkedIn "Experience" to Twitter "User Mentions" (@tags), assuming professionals mention their employers in tweets.
- Achievements (SoftTFIDF): A sophisticated metric that captures "similar" rather than identical tokens, linking job titles to bio descriptions (e.g., "Editor" vs. "Journalist").
2. The Hybrid Classifier: DTNB
The heart of the approach is the DTNB (Decision Table Naive Bayes) classifier. It combines the rule-based strengths of Decision Tables with the probabilistic robustness of Naive Bayes, assigned the maximum probability to determine a "Match" even under heavy data imbalance.
Experiments & Results: Professional Signals Win
The authors tested their approach on 262 "target users." A critical finding was the value of achievements. On LinkedIn, the "Achievements" metric alone achieved 87% recall, highlighting that professional context is more stable across platforms than names or locations.
Fig 3: ROC curves showing DTNB outperforming Decision Trees and Naive Bayes in LinkedIn identification.
Key Quantitative Results:
| Task | Precision | Recall | F-measure |
|---|---|---|---|
| LinkedIn Identification | 94.98% | 94.27% | 94.62% |
| Twitter Identification | 90.73% | 90.73% | 90.73% |
The study also addressed missing values—a chronic issue in social data. They found that replacing missing scores with the average similarity score of other available fields outperformed setting a default constant (0.5) or using the median.
Critical Analysis & Conclusion
Takeaway
The research proves that cross-network identity is not just about the name; it’s about the semantic overlap of professional persona. By aligning unstructured text (mentions) with structured data (affiliations), we can bypass the noise of social media.
Limitations
- Search Engine Dependency: The model relies on the top 25 results from native search engines. If the engine fails to retrieve the profile at all, the classifier can't find it.
- Static Features: The model doesn't account for temporal changes in profiles (e.g., changing jobs).
Future Work
The authors suggest integrating SMOTE for better imbalance handling and exploring the identification of fake or compromised accounts—a vital next step in ensuring social media integrity.
