Decoding Gender in Vietnamese Names: A Deep Learning Benchmark

Gender Prediction Based on Vietnamese Names with Machine Learning Techniques

2020-12-18
Huy Quoc To, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen, Anh Gia-Tuan Nguyen
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces UIT-ViNames, a novel dataset of over 26,000 annotated Vietnamese full names for gender prediction. The study benchmarks six traditional machine learning algorithms and a Long Short-Term Memory (LSTM) model with fastText embeddings, achieving a SOTA F1-score of 96%.

Executive Summary

TL;DR: Researchers from the University of Information Technology (VNU-HCM) have developed UIT-ViNames, a massive dataset of 26,850 Vietnamese names, and proved that LSTM models combined with fastText embeddings can predict biological gender with an impressive 96% F1-score. The study uniquely highlights that in the Vietnamese context, the "Middle Name" is the secret sauce for high-accuracy classification.

Positioning: This work fills a significant gap in regional NLP. While most global systems struggle with non-Western name structures, this paper provides both the data and the optimized architectural blueprint specifically for the Vietnamese linguistic landscape.

The "Nguyen" Problem: Why Surnames Don't Matter

In many Western cultures, a surname can sometimes hint at heritage, but in Vietnam, the surname distribution is extremely skewed. Approximately 40% of the population shares the surname "Nguyá»…n".

The authors' initial analysis (Figure 1) confirms a critical intuition: the distribution of surnames between males and females is virtually identical. Using a surname to predict gender in Vietnam is mathematically equivalent to a random guess.

Distribution of surnames Figure 1: Comparison showing that common Vietnamese surnames (Nguyễn, Trần, Lê) provide zero discriminatory power for gender.

Methodology: Beyond Simple Dictionaries

The study compared traditional "Bag-of-Words" approaches with sequential Deep Learning.

1. Traditional Baselines

The team tested SVM, Multinomial Naive Bayes, Bernoulli NB, Logistic Regression, and Random Forest. Using TF-IDF and Count Vectorization, these models performed admirably (around 94-95% accuracy), but struggled with "unisex" names.

2. The Winning Combo: LSTM + fastText

The core of the successful approach was a Long Short-Term Memory (LSTM) network. Unlike traditional models, LSTMs can maintain the "memory" of name order. By using fastText embeddings (300-dimension), the model captures the morphological relationships between names even when they haven't been seen in the training set.

Name Component Distribution Figure 3: The distribution of common first names highlights gender-specific clusters (e.g., "Thị" for females, "Văn" for males).

Experimental Insights: The Power of the Middle Name

The most technical revelation of the paper comes from the Ablation Study. The researchers systematically stripped away parts of the name to see what happened to the accuracy:

Name CombinationSVM (Avg F1)LSTM (Avg F1)
Family Name Only39.60%38.23%
First Name Only87.04%80.02%
Middle + First Name95.28%95.89%

Key Insight: Standalone first names (like "Anh" or "Tú") are often gender-neutral. Adding the Middle Name acts as a disambiguator. For example, "Tuấn Anh" is almost certainly male, while "Tú Anh" is frequently female.

Critical Analysis & Conclusion

The "Unseen" Challenge

Despite the high accuracy, the model still fails on rare or modern naming conventions. For instance, the name "Lâm Anh" (typically female) was misclassified when assigned to a male in the dataset. These "edge cases" represent the cultural shift in Vietnam toward more unique, non-traditional names.

Takeaways

  1. Architecture Matters: For short-text sequences like names, LSTMs provide the necessary sequential context that Naive Bayes lacks.
  2. Data Sparsity: Surnames in Vietnamese are "noise"—removing them simplifies the model and potentially improves generalization.
  3. Future Path: The authors suggest moving toward Transformer-based models (BERT) to see if bidirectional context can push that 96% closer to 100%.

This research provides a robust API-ready solution for automated form-filling, coreference resolution, and demographic analysis in the Vietnamese digital ecosystem.

Find Similar Papers

Try Our Examples

  • Examine recent SOTA papers on gender prediction from names in other Southeast Asian languages with similar naming structures.
  • Which paper originally established the fastText word vector approach for 157 languages, and how does it handle the specific diacritics of the Vietnamese alphabet?
  • Investigate the use of BERT-based transfer learning models specifically fine-tuned for Vietnamese text classification and compare their performance to LSTM for short-text tasks.
Contents
Decoding Gender in Vietnamese Names: A Deep Learning Benchmark
1. Executive Summary
2. The "Nguyen" Problem: Why Surnames Don't Matter
3. Methodology: Beyond Simple Dictionaries
3.1. 1. Traditional Baselines
3.2. 2. The Winning Combo: LSTM + fastText
4. Experimental Insights: The Power of the Middle Name
5. Critical Analysis & Conclusion
5.1. The "Unseen" Challenge
5.2. Takeaways