Age and Gender Estimation: Is Deeper Always Better? Finding the Architectural Sweet Spot

Age and Gender Estimation using Optimised Deep Networks

2019-09-03
Wade Downton, Hima Vadapalli
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores the optimization of Convolutional Neural Networks (CNNs) for age and gender estimation using the Adience dataset. It specifically investigates the impact of various activation functions, including the Swish function, and model depth on classification accuracy, ultimately achieving gender estimation performance (79%) comparable to existing benchmarks with a significantly reduced parameter count.

Executive Summary

TL;DR: This study investigates how to optimize CNNs for facial attribute estimation by balancing depth and activation function choice. The researchers found that for tasks like gender classification, a shallower, leaner network can match SOTA performance (approx. 79-81%) with significantly fewer parameters, whereas age estimation remains a bottleneck due to its complex feature nature.

The work stands as a pragmatic optimization study, shifting the focus from "scaling up" (more layers) to "scaling right" (better parameters and activation functions).

Problem and Motivation: The Wall of Over-fitting

In the realm of Computer Vision, the reflex is often to add more layers for better features. However, age and gender estimation from "in-the-wild" images (like those in the Adience dataset) present a unique challenge: over-fitting. Unlike object detection, facial aging features are subtle and highly sensitive to lighting, angle, and resolution.

The authors noted that deep, fully-connected architectures often lead to "feature drifting," where the model learns noise rather than the biological markers of age. This motivated an exploration into whether minimizing complexity could actually enhance reliability.

Methodology: The Core of Optimization

The researchers used a base architecture of 3 Convolutional stages and 3 Fully-Connected stages. Their key experimental variables were:

  1. Network Depth: Systematically adding Conv or FC layers to observe the "breaking point" of convergence.
  2. Activation Functions: Comparing the standard ReLU against Linear, ELU, tanh, and the then-novel Swish ().

Full Network Architecture

The Role of Swish

The authors specifically tested Google's Swish function. Because Swish is non-monotonic and unbounded above, it avoids the vanishing gradient problem of tanh while offering a smoother landscape than ReLU.

Experimental Insights: Less is More

The results revealed a surprising trend. For both gender and age, the Base Model far outperformed deeper variants.

1. The Depth Penalty

Adding just one or two more layers caused a notable decrease in the convergence rate. For age estimation, adding convolutional layers actually caused the model to diverge (the error increased over time). This suggests that the extra capacity was being used to memorize noise.

Loss Term Comparison Figure: The Base Model (top lines) showing much faster loss convergence compared to augmented depths.

2. The Activation Battle

  • ReLU & Swish: Tied for the best performance.
  • Linear: Slightly worse, lacking the non-linearity needed to map complex facial features.
  • tanh & ELU: Struggled with convergence, likely due to saturation (tanh) or insufficient training epochs for the more complex ELU landscape.

3. Age vs. Gender: A Complexity Gap

While gender classification hit a respectable 79%, age estimation hovered around 40%. The confusion matrices indicate that while models are great at identifying infants (0-2 years), they often confuse middle-age brackets (25-32 vs 38-43).

Confusion Matrices

Deep Insights & Conclusion

Takeaway: This paper provides a sobering reminder that Inductive Bias (the assumptions we build into our model) is more important than raw size. For facial attributes:

  • Unbounded activations (ReLU/Swish) are essential.
  • Shallow networks act as a natural regularizer against over-fitting in low-quality datasets.
  • Data Imbalance remains a primary enemy—even class weighting cannot fully compensate for a lack of representative samples in specific age groups (e.g., the 48-53 bracket).

Limitations: The study was limited to 10 epochs. Future work should investigate if deeper models like ResNets or RoR could eventually surpass these shallow baselines if trained for hundreds of epochs with stronger regularization.

Future Outlook: The parity between ReLU and Swish suggests that for mobile or edge applications, the simpler ReLU is likely sufficient, whereas Swish might offer marginal gains only in extremely high-resolution tasks.

Find Similar Papers

Try Our Examples

  • Find recent papers that address age estimation as a regression problem rather than a multi-class classification task to improve accuracy.
  • Which original research first introduced the Swish activation function, and how did its self-gating mechanism statistically outperform ReLU in deep vision tasks?
  • Explore current SOTA methods for handling extreme class imbalance in facial attribute datasets beyond simple class weighting or oversampling.
Contents
Age and Gender Estimation: Is Deeper Always Better? Finding the Architectural Sweet Spot
1. Executive Summary
2. Problem and Motivation: The Wall of Over-fitting
3. Methodology: The Core of Optimization
3.1. The Role of Swish
4. Experimental Insights: Less is More
4.1. 1. The Depth Penalty
4.2. 2. The Activation Battle
4.3. 3. Age vs. Gender: A Complexity Gap
5. Deep Insights & Conclusion