Gender Identification from Bengali Names: A Comparative Supervised Learning Approach

Performance Measurement of Multiple Supervised Learning Algorithms for Gender Identification from Bengali Names

2021-07-06
Labannya Saha, Rakib Md. Azhar Uddin, Sorna Saha
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a comparative study of seven supervised machine learning algorithms for identifying gender from Bengali names. By employing Natural Language Processing (NLP) techniques like character-level N-grams and Count Vectorizer, the researchers achieved a peak accuracy of 84% using the Multinomial Naive Bayes classifier.

TL;DR

This study addresses the complex task of predicting gender from Bengali names using NLP and Machine Learning. By benchmarking seven distinct algorithms, the researchers identified Multinomial Naive Bayes as the most effective model, achieving an 84% accuracy. This work paves the way for automated demographic analysis in the Bengali-speaking world, which comprises over 230 million people globally.

Background & Motivation

Identifying gender from names is a cornerstone of modern NLP applications, ranging from personalized recommendation systems to targeted text summarization. However, Bengali presents unique hurdles:

  • Linguistic Roots: Names are often derived from Arabic, Persian, or Sanskrit, each with different gender markers.
  • Morphological Complexity: Bengali script uses "Kar" (vowel signs) and "Jukto-Borno" (clusters) that carry subtle gender cues.
  • Unisex Names: Many names are culturally valid for both males and females, requiring models that can capture nuanced sub-word patterns.

Methodology: The Core Architecture

The researchers treated gender identification as a document classification problem at the character level. The workflow is structured as follows:

  1. Data Pre-processing: Removal of duplicates and handling missing values from a curated set of 1,610 names.
  2. Feature Extraction (Bigrams): Instead of treating names as whole words, the authors broke them into substrings (e.g., 'লাবন্য' becomes ['লা', 'াাব', 'বন্', 'ন্্', 'া্য']). This allows the model to learn that certain suffixes or character combinations are more frequent in specific genders.
  3. Vectorization: Using Count Vectorizer to convert text into numeric frequency vectors.
  4. Scaling: Applying Min-Max Normalization to ensure all features reside in a [0, 1] range.

Architecture of the Methodology

SOTA Comparison and Performance

The study pitted classical algorithms against each other. Probabilistic and linear models generally outperformed tree-based ensemble methods:

  • Naive Bayes (84%): The winner. Its ability to calculate conditional probabilities based on independent features (character bigrams) proved superior for this specific sparse data.
  • Logistic Regression (83%): Performed remarkably well, acting as a strong baseline for binary gender classification.
  • Ensemble Methods: Random Forest (82%) and AdaBoost (77%) were slightly less effective, likely due to the limited size of the dataset (1.6k samples), where simple probabilistic models often generalize better.

Performance Comparison Graph

Detailed Metrics

AlgorithmAccuracyPrecisionRecallF1-Score
Naive Bayes0.840.830.680.84
Logistic Regression0.830.830.670.83
Random Forest0.820.820.640.82
SVM0.820.820.650.82

Critical Insight & Future Outlook

The primary takeaway is that sub-word features (N-grams) are critical for Bengali. While an 84% accuracy is a strong start, the authors acknowledge a limitation: the Sample Size. In the era of Deep Learning, a dataset of 1,610 names is relatively small.

Future Directions:

  • Data Scaling: Expanding the dataset across various Bengali ethnicities (West Bengal vs. Bangladesh) will likely improve generalization.
  • Neural Networks: Transitioning to Bi-LSTMs or BERT-based architectures could capture long-range dependencies in longer names that N-grams might miss.
  • Deployment: The authors aim to release a Gender API, which could become a vital tool for researchers working on Bengali social media sentiment analysis and demographic profiling.

Conclusion: This paper establishes a solid baseline for Bengali gender identification, proving that even "traditional" machine learning algorithms like Naive Bayes can remain competitive when paired with thoughtful feature engineering.

Find Similar Papers

Try Our Examples

  • Find recent papers from 2023-2025 that use Deep Learning architectures like LSTMs or Transformers for gender identification specifically in South Asian or Indic languages.
  • Which paper first established the use of character-level N-grams for name-based gender classification, and how do modern Bengali-specific methods improve upon those feature extraction techniques?
  • Explore how gender identification algorithms from names are being integrated into large-scale social media analytics or customer relationship management (CRM) systems in the context of multilingual populations.
Contents
Gender Identification from Bengali Names: A Comparative Supervised Learning Approach
1. TL;DR
2. Background & Motivation
3. Methodology: The Core Architecture
4. SOTA Comparison and Performance
4.1. Detailed Metrics
5. Critical Insight & Future Outlook