VGG-16 vs. MobileNet: Benchmarking Gender Classification for Asian Demographics

Gender Classification Based on Asian Faces using Deep Learning

2019-10-01
Tiagrajah V. Janahiraman, Prasantth Subramaniam
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores gender classification specifically targeting Asian faces by benchmarking three Deep Learning architectures: VGG-16, ResNet-50, and MobileNet. Leveraging transfer learning with ImageNet pre-trained weights, the study identifies VGG-16 as the most effective model for this demographic, achieving a 100% training accuracy and an 88% recognition rate on real-world test images.

TL;DR

This research addresses the "representation gap" in facial recognition by focusing on Asian demographics. By benchmarking three industry-standard models—VGG-16, ResNet-50, and MobileNet—the authors demonstrate that while lightweight models are efficient, VGG-16 remains the superior choice for accuracy, achieving an 88% recognition rate on a custom Asian face dataset compared to a disappointing 49% from MobileNet.

Context & Motivation

Despite the explosion of AI-driven surveillance and human-computer interaction (HCI) tools, gender classification accuracy remains inconsistent across different ethnicities. Most standard models are trained on datasets with a heavy Caucasian bias (like LFW or Adience).

The authors identify two major hurdles:

  1. Lack of Diversity: Existing SOTA models often fail on "real-world" Asian faces due to data distribution shifts.
  2. Localized Data Scarcity: There is a lack of publicly available, high-quality facial databases specifically featuring Malaysian and wider Asian demographics.

Methodology: The Transfer Learning Approach

The study utilizes Transfer Learning, taking models pre-trained on the massive ImageNet dataset and fine-tuning them on a specialized niche.

The Pipeline

  1. Pre-processing: Faces are detected and extracted using the Haar Cascade algorithm via OpenCV.
  2. Architecture Selection:
    • VGG-16: Known for its simplicity and depth (138 million parameters).
    • ResNet-50: Utilizes "skip connections" to mitigate vanishing gradients.
    • MobileNet: Designed for mobile efficiency using depth-wise separable convolutions.
  3. Training: Fixed at 100 epochs with a learning rate of 0.001 using the SGD optimizer.

System Architecture Fig 1: The proposed classification workflow from image input to probability output.

Experiments & Performance Analysis

The authors built the U10 Face Database, consisting of 1,000 images (50/50 male-female split). The results highlight a stark contrast between model complexity and performance on this specific task.

Training Convergence

All models showed high capability on the training set (all >99%), but the real test was the Recognition Rate on unseen faces.

ModelTraining AccuracyActual Recognition RateParameters
VGG-16100%88%138.3M
ResNet-5099.9%85%25.6M
MobileNet99.8%49%4.2M

The "MobileNet Failure" Insight

The most striking result is MobileNet's 49% accuracy—essentially no better than a coin flip for a binary class. This suggests that the depth-wise separable convolutions, while excellent for saving power, may discard the fine-grained spatial textures required to distinguish gender in Asian faces when the dataset size is limited.

Training Loss Curves Fig 2: Training loss visualization showing rapid convergence across all models.

Critical Analysis & Takeaways

The paper confirms that for high-stakes biometric identification, bigger is often better when it comes to feature extraction power. VGG-16's dense parameterization allows it to capture subtle facial markers that ResNet and MobileNet missed.

Limitations:

  • Dataset Size: 1,000 images is relatively small for deep learning; this likely contributed to the high training accuracy (overfitting) and the struggle of the lightweight MobileNet.
  • Demographic Granularity: While focusing on "Asian faces," further sub-categorization (East Asian vs. South Asian) could yield even more specific insights into model bias.

Future Work: The authors suggest exploring InceptionResNetV2 and DenseNet to see if feature concatenation or multi-scale processing can push the 88% recognition rate closer to 95%+.

Conclusion

This study serves as a vital benchmark for localized AI applications. It proves that while we strive for "mobile-first" AI, we cannot sacrifice the depth required to handle demographic diversity. For developers building HCI systems in Asia, VGG-based backbones remain the gold standard for reliable gender classification.

Find Similar Papers

Try Our Examples

  • Find recent papers from 2023-2025 that address gender classification bias specifically in Southeast Asian and East Asian facial datasets.
  • Which paper first introduced Depthwise Separable Convolutions in MobileNet, and how does its reduced parameter count specifically impact feature representational power in facial analysis?
  • Search for studies that apply Vision Transformers (ViT) to the same Asian face gender classification task to see if self-attention outperforms the CNN localized feature extraction mentioned in this paper.
Contents
VGG-16 vs. MobileNet: Benchmarking Gender Classification for Asian Demographics
1. TL;DR
2. Context & Motivation
3. Methodology: The Transfer Learning Approach
3.1. The Pipeline
4. Experiments & Performance Analysis
4.1. Training Convergence
4.2. The "MobileNet Failure" Insight
5. Critical Analysis & Takeaways
6. Conclusion