Bridging the Recruitment Gap: A Bottom-Up Multilingual Skills Knowledge Base

Bridge the terminology gap between recruiters and candidates: A multilingual skills base built from social media and linked data

2016-08-01
Emmanuel Malherbe, Marie-Aude Aufaure
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a novel bottom-up approach to building a multilingual skills knowledge base by analyzing terminology used by candidates on professional social networks. By integrating Social Media data with Linked Open Data (DBpedia) and technical Q&A tags (StackOverflow), the authors developed a system that outperforms traditional top-down taxonomies in skills extraction and normalization.

TL;DR

Recruiters and candidates often speak different languages—not just literally, but terminologically. This paper introduces a method to build a massive, multilingual skills base by mining professional social networks and linking them to DBpedia and StackOverflow. The result? A system that captures "real-world" skills 5x more effectively than traditional HR standards like ESCO or O*Net.

The Problem: The Ivory Tower of HR Taxonomies

In the world of E-Recruitment, the "Skill" is the fundamental unit of value. However, current systems struggle with Entity Linking. If a candidate writes "Dynamic Scaling in AWS" and a job ad asks for "Cloud Infrastructure Management," traditional keyword-matching fails.

Existing taxonomies (ESCO, O*Net) are built "top-down" by committees. They are:

  1. Outdated: They miss emerging tech (e.g., "Prompt Engineering").
  2. Formalistic: They use academic terms that neither the hiring manager nor the candidate actually types.
  3. Language-Isolated: Mapping "Soudage" (FR) to "Welding" (EN) usually requires expensive manual translation.

Methodology: Let the Data Speak

The authors propose a bottom-up approach. Instead of defining skills first, they look at what 4.3 million candidates say they can do.

1. Data Collection & Filtering

The system aspirates data from over 120 sources (Indeed, XING, Viadeo). It identifies expressions that appear in more than 0.01% of profiles as potential "Skill Entities."

2. Linking to Global Knowledge

To turn a "flat string" into a "rich entity," the system maps these terms to:

  • DBpedia: Provides cross-language "sameAs" links, descriptions, and categories.
  • StackOverflow Tags: Crucial for the fast-moving software domain where DBpedia might lag.

System Architecture The workflow from raw social media data to a structured, merged knowledge base.

3. Merging and Multi-linguality

By using DBpedia's inter-language links, a skill identified in a French profile becomes automatically linked to its English equivalent.

Experiments: Performance Over Precision

The researchers compared their "Final Base" against industry standards using two main metrics: Normalization Coverage (how many terms in a profile can we identify?) and Extraction Precision (how accurate are the tags provided for a job ad?).

MetricExisting (ESCO)Our Final Base
Coverage (FR Profiles)13.9%76.9%
Coverage (EN Profiles)16.1%73.8%
Precision (Job Ads)73.9%81.0%

The jump from ~15% to ~75% coverage is massive. It suggests that traditional HR bases miss nearly 80% of the relevant skills actually mentioned by professionals online.

Experimental Results

Real-World Impact: The SmartSearch Project

This isn't just a theoretical exercise. The system was implemented for SAP subsidiary Multiposting. It allows for "Market Analysis" dashboards that can track "Skill Mismatch" (the delta between candidate supply and recruiter demand) in real-time.

Market Analysis Dashboard Example: Analyzing 'Welding' skills across languages and companies.

Critical Insight & Conclusion

The genius of this paper lies in its Inductive Bias: it assumes that the crowd (social media) is more accurate at defining a domain than a central authority.

Limitations: While the coverage is excellent, the system still relies on n-gram matching. This can lead to "false positives" if a skill name is a common word (e.g., "Sketch" or "Python" in a non-coding context). Future iterations would benefit from Dependency Parsing or Modern Transformer-based NER (Named Entity Recognition) to understand context.

The Takeaway: In e-recruitment, connectivity is king. By linking "Bottom-Up" social data with "Top-Down" Linked Data, we get a system that is both broad enough to cover the market and deep enough to provide meaningful analytics.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Large Language Models (LLMs) to bridge the terminology gap in e-recruitment compared to the graph-based entity linking approach used here.
  • Which studies first established the "bottom-up" methodology for ontology maturing in professional networks, and how has this evolved since the 2016 SmartSearch project?
  • Explore how multilingual skills extraction techniques have been applied to multi-modal recruitment data, such as video interviews or LinkedIn portfolio scraping.
Contents
Bridging the Recruitment Gap: A Bottom-Up Multilingual Skills Knowledge Base
1. TL;DR
2. The Problem: The Ivory Tower of HR Taxonomies
3. Methodology: Let the Data Speak
3.1. 1. Data Collection & Filtering
3.2. 2. Linking to Global Knowledge
3.3. 3. Merging and Multi-linguality
4. Experiments: Performance Over Precision
5. Real-World Impact: The SmartSearch Project
6. Critical Insight & Conclusion