LinkedIn Skills: Mapping the Global Professional Genome via Folksonomy and Inference
LinkedIn skills: large-scale topic extraction and inference
This paper presents LinkedIn's large-scale "Skills and Expertise" system, detailing the creation of a massive professional folksonomy and a recommendation engine. By combining NLP, clustering, and Naive Bayes inference, the system automates skill tagging for hundreds of millions of professional profiles.
TL;DR
LinkedIn revolutionized how professional expertise is captured by moving from a static, expert-curated taxonomy to a data-driven folksonomy. By extracting entities from millions of profiles, disambiguating them via clustering, and leveraging a Naive Bayes recommender system, they boosted skill-tagging conversion by over 1,200%.
Background & Motivation: The Problem of "The Blank Box"
How do you categorize the skills of 300M+ professionals ranging from ballet dancers to nuclear engineers? LinkedIn's initial attempts were thwarted by two main issues:
- Taxonomy Scalability: No existing list was comprehensive enough to cover the "long-tail" of specialized professional niches.
- The Interaction Gap: Simply providing a search box (Type-ahead) led to stagnant profiles. Users are much more likely to confirm a suggestion than to recall and type one from scratch.
Authors identified that professional identity is often "hidden in plain sight" within free-text profile sections. The challenge was converting this messy text into standardized, deduplicated, and disambiguated "entities."
Methodology: The Three Pillars of Folksonomy
1. Discovery & Entity Extraction
The team focused on the "Specialties" section of profiles, identifying comma-separated lists using a specific punctuation frequency threshold (). This allowed them to filter prose from keywords, yielding 150,000 candidate phrases.
2. Disambiguation: The "Organ" Problem
A skill like "Organ" means something different to a Cardiologist than to a Jazz Musician. To solve this, the authors used Co-occurrence Clustering:
- Jaccard Similarity: Measured how often two phrases (e.g., "Organ" and "Surgery") appeared together.
- SVD & KMeans: Applied dimensionality reduction to the similarity matrix to identify distinct "senses" of a word.
- Industry Labeling: Attached the most common member industry to each cluster to provide context.
Figure 1: The Skills & Expertise section UI, the final output of the pipeline.
3. Deduplication via Crowdsourcing
Members use synonyms like "Java Programming" and "Java Development." The authors cleverly mapped these to Wikipedia entities using Amazon Mechanical Turk. If two different phrases mapped to the same Wikipedia URL, they were merged into a single standardized skill.
Inference: Solving the Cold Start
To suggest skills for users who hadn't listed any, LinkedIn built a Naive Bayes Classifier. Instead of looking at past skills, it looks at "Collaborative Attributes": If your peers at "Google" with the title "Software Engineer" all have the skill "Distributed Systems," the system infers that you likely do too.
Experimental Results: Performance and UX
The most striking result wasn't just accuracy, but User Behavior shift.
Figure 2: Type-ahead (4% conversion) vs. Recommendations (49% conversion).
- Conversion: By switching to a 10-item recommendation list, conversion jumped from 4% to 49%.
- AUC Performance: The inference engine achieved an AUC of 0.77. Technical skills (e.g., "Hadoop") performed better because they have "tighter" co-occurrence patterns compared to "soft skills" like "Teamwork."
Critical Insight & Conclusion
The genius of the LinkedIn Skills system lies in its recognition that Identity is Social. By using "Endorsements" as a social gesture and "Inference" as a psychological nudge, they turned a data-entry chore into a core part of professional reputation.
Limitations: The system relies heavily on professional homophily (the idea that people with similar titles have similar skills). This can create a "filter bubble" where rare or cross-disciplinary skills are harder to infer.
Future Work: Transitioning from Naive Bayes to more complex Machine Learned models using real-world user feedback (accept/reject signals) will likely close the gap on "soft skill" inference accuracy.
