Deconstructing the Job Title: A Structural Approach to Semantic Similarity

Similarity Computation Exploiting the Semantic and Syntactic Inherent Structure Among Job Titles

2017-01-01
Sarthak Ahuja, Joydeep Mondal, Sudhanshu Shekhar Singh, David Glenn George
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a specialized "Parts Of Title" (POT) tagging framework to compute similarity between job titles by decomposing them into Domain, Function, and Attribute components. It utilizes Machine Learning classifiers and WordNet-based semantic matching to outperform standard document similarity methods in recruitment analytics.

TL;DR

In the world of HR Tech and recruitment analytics, knowing if two jobs are "the same" is a billion-dollar problem. This paper from IBM Research proposes a "Parts Of Title" (POT) framework that breaks job titles into Domain, Function, and Attribute. By treating title matching as a structured assignment problem rather than simple keyword matching, the authors achieve over 90% accuracy in identifying the core components of a job, significantly improving the quality of job clustering.

Context & Positioning

Most search engines treat text as a "bag of words." If you search for "Senior Software Engineer," a standard algorithm might give equal weight to "Senior" and "Software." However, in recruiting, the "Software" part (Domain) is non-negotiable, while "Senior" (Attribute) is a modifier. This paper moves away from flat document similarity (TF-IDF/LSA) and creates a linguistic ontology specifically for the labor market.

The Core Insight: The Trinity of a Job Title

The authors argue that every professional title is composed of three inherent dimensions:

  1. Domain: The industry or field (e.g., Software, Electrical, Dental).
  2. Function: The actual role or action (e.g., Engineer, Manager, Consultant).
  3. Attribute: The seniority or level (e.g., Senior, Junior, Lead).

The challenge? The same word can change meaning based on its position. "Assistant" in "Assistant Manager" is an Attribute, but in "Lab Assistant," it is the Function.

Methodology: How POT Works

The system operates in a two-stage pipeline: Extraction and Matching.

1. Sequential Tagging (The Architecture)

Instead of labeling words in isolation, the authors found that identifying one component helps identify the others. They tested 16 different dependency permutations. The winner? A linear flow: Function → Attribute → Domain.

Model Architecture Fig 1: The Training Phase where feature vectors (Position, Suffix, POS Tags) are fed into an ensemble to find the best dependency arrangement.

2. The Hungarian Matchmaker

Once words are tagged, you can't just average their scores. If Title A has two "Domain" words and Title B has one, how do you pair them? The authors framed this as an Assignment Problem. Using the Hungarian Method, the system finds the optimal one-to-one mapping between keywords of the same context to maximize the total semantic score.

Experiments and Results

The model was tested using the IBM Talent Framework. By comparing "Intra-family" (jobs in the same category) vs. "Inter-family" (jobs in different categories) similarity, the researchers proved their method captures the "essence" of a job better than standard WordNet averages.

Experimental Results Table 1: Comparison of different model arrangements. Notice how the 'f -> a -> d' arrangement yields the highest validation accuracy (90.76%).

Key Metrics:

  • Attribute Accuracy: 93.43% (The easiest to identify due to common suffixes like -sr/jr).
  • Function Accuracy: 87.01%.
  • Domain Accuracy: 78.04% (The most complex due to the vast vocabulary of industries).

Critical Analysis & Conclusion

Why this matters

For hiring platforms, this provides Explainable AI. If a system recommends a job, it can now say: "This matched because the Domain (Software) and Function (Engineer) are identical, even though the Attribute (Level) is higher." This is a massive leap over "black-box" vector similarities.

Limitations

The system relies heavily on WordNet and hard-coded abbreviation lists. In a rapidly evolving economy where titles like "Cloud Evangelist" or "Happiness Officer" emerge, a purely taxonomic approach might struggle. Future work suggests integrating Word2Vec or Compound Noun Semantics to better handle these creative titles.

Final Takeaway

Structure beats raw data. By understanding the "Syntactic Inherent Structure" of a specific domain (Recruitment), we can build lighter, more accurate, and more explainable models than by simply throwing more parameters at the problem.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Deep Learning-based Named Entity Recognition (NER) specifically for job title parsing and attribute extraction.
  • Find the original paper detailing the Wu-Palmer (WUP) Similarity measure and its effectiveness in short-text semantic matching compared to BERT embeddings.
  • Explore how the Hungarian Method is applied in modern natural language processing tasks beyond synonym matching, such as in graph-based text summarization.
Contents
Deconstructing the Job Title: A Structural Approach to Semantic Similarity
1. TL;DR
2. Context & Positioning
3. The Core Insight: The Trinity of a Job Title
4. Methodology: How POT Works
4.1. 1. Sequential Tagging (The Architecture)
4.2. 2. The Hungarian Matchmaker
5. Experiments and Results
5.1. Key Metrics:
6. Critical Analysis & Conclusion
6.1. Why this matters
6.2. Limitations
6.3. Final Takeaway