What’s in a Name? Decoding Gender through Character-Based Machine Learning

What’s in a name? – gender classification of names with character based machine learning models

2021-05-12
Yifan Hu, Changwei Hu, Thanh Tran, Tejaswi Kasturi, Elizabeth Joseph, Matt Gillingham
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces character-based machine learning models for gender classification of names, leveraging a massive dataset of 100M+ users (Yahoo Data). It demonstrates that name strings—both first and last—encode sufficient gender information to achieve high classification accuracy, significantly outperforming content-based behavior models.

TL;DR

Gender information is becoming increasingly scarce as users opt out of demographic disclosure. This paper demonstrates that we don't need user activity logs to guess gender; the composition of the name itself is a high-fidelity signal. By using character-based models on a dataset of 100 million users, the authors achieved an AUC of 0.94, proving that the way a name is spelled contains deep-seated societal "gender coding."

The "Sparse Signal" Problem

Traditional demographic inference often relies on Content-Based Models—analyzing what a user clicks or reads. However, this method has two fatal flaws:

  1. Sparsity: New or inactive users have no history.
  2. Noise: Content categories (e.g., "Sports" or "Cooking") are increasingly gender-neutral, leading to modest AUCs around 0.80.

The authors pivot to names, but not via simple lookups. They tackle the "Long Tail" problem where 72% of names in their dataset appear only once. This cardinality makes traditional word embeddings (word2vec) fail. Their solution? Character-level analysis.

Methodology: From N-Grams to Transformers

The researchers tested an array of models, ranging from "classic" ML to "SOTA" deep learning:

1. NBLR (Class Scaled Logistic Regression)

Surprisingly, the star of the show was a linear model. By extracting character n-grams (up to 7-grams) and applying a log-ratio scaling similar to NBSVM, it captured prefixes and suffixes that are strong gender indicators (e.g., names ending in "-a" in English are 34.7% female vs 1% male).

2. Dual-LSTM (The Cultural Context Bridge)

A significant contribution is the DLSTM architecture. First names like "Andrea" are unisex or gendered differently across cultures (Male in Italy, Female in the US). By processing the last name in a parallel LSTM arm, the model learns the cultural/ethnic context needed to disambiguate these cases.

Model Architecture Figure 1: The Dual-LSTM (DLSTM) architecture utilizing both first and last names to resolve unisex ambiguities.

Experiments & SOTA Results

The models were trained on 21M unique first names from Yahoo Data and validated against the US Social Security Administration (SSA) dataset.

  • The Linear Surprise: NBLR achieved an AUC of 0.940, virtually identical to the much heavier Char-BERT (0.935) and LSTM (0.938).
  • Unisex Breakthrough: On names like "Toni" and "Andrea," using the last name via DLSTM boosted AUC from near-random (0.50) to ~0.85.
  • Generalizability: Models trained on the multilingual Yahoo dataset generalized perfectly to the English-only SSA data, but not vice-versa, highlighting the importance of high-volume, diverse training data.

Performance Visual Figure 2: Accuracy comparison of DLSTM vs. LSTM. The gap widens significantly for unisex names (middle of the X-axis), proving that last names provide crucial context.

Deep Insight: Why did the Linear Model Win?

In many NLP tasks, Transformer-based models like BERT dominate. Why did a simple Logistic Regression (NBLR) hold its own here?

  1. Structural Simplicity: Gender signals in names are often localized at the start or end of the string (suffixes/prefixes). N-grams explicitly capture these without needing the complex self-attention layers of a Transformer.
  2. Data Volume: With 100M+ users, the feature space for 7-grams becomes dense enough that linear separation becomes highly effective.

Critical Analysis & Future Work

While the results are impressive, the study acknowledges a major limitation: Name Validity. The model assumes users register with "real" names. If a user registers as "Mickey Mouse," the model will blindly predict a gender based on those characters.

The ultimate potential lies in Model Fusion. The authors show that combining name-based predictions with content behavior via an XGBoost ensemble yields the highest possible accuracy (AUC 0.978). For the industry, this represents a robust pipeline for unintentional bias intervention in AI systems.

Summary Takeaway

If you're building a demographic inference engine, don't ignore the characters. A well-engineered linear model on name strings is often more efficient and just as accurate as a deep neural network, especially when dealing with the vast, multilingual long-tail of global names.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize character-based transformers or state-space models for demographic inference beyond gender, such as age or ethnicity prediction.
  • Which paper first established the concept of "ethnic and gender homophily" in digital contact lists, and how has this theory evolved with large-scale graph neural networks?
  • Investigate if character-based name classification models have been successfully applied to non-Latin scripts (e.g., Arabic or Devanagari) and what specific inductive biases are required for those languages.
Contents
What’s in a Name? Decoding Gender through Character-Based Machine Learning
1. TL;DR
2. The "Sparse Signal" Problem
3. Methodology: From N-Grams to Transformers
3.1. 1. NBLR (Class Scaled Logistic Regression)
3.2. 2. Dual-LSTM (The Cultural Context Bridge)
4. Experiments & SOTA Results
5. Deep Insight: Why did the Linear Model Win?
6. Critical Analysis & Future Work
7. Summary Takeaway