Beyond Voice Pitch: Improving Gender Identification with Multi-Task Age Context
Multi-task learning DNN to improve gender identification from speech leveraging age information of the speaker
The paper introduces a Multi-Task Learning (MTL) Deep Neural Network (DNN) that improves gender identification from speech by using the speaker's age as an auxiliary target. The system employs an end-to-end architecture featuring a raw waveform front-end with 1D-Convolutional layers and LSTMP (Long Short Term Memory with Recurrent Projection) layers for temporal modeling.
TL;DR
Researchers have developed a new Deep Neural Network (DNN) architecture that solves a classic problem in speech analysis: identifying gender across all ages. By teaching a model to recognize both age and gender simultaneously from raw audio waves (avoiding traditional MFCCs), they achieved a massive 29.9% relative error reduction compared to traditional methods on diverse real-world datasets.
Contextualizing the Challenge
Gender identification might seem like a "solved" problem in AI, with many models hitting 98%+ accuracy on adult speech. However, these models often fail when encountering children or seniors. In kids, the physiological differences in vocal tracts haven't fully diverged yet, and in seniors, hormonal and physical changes can blur vocal gender lines.
The authors of this paper realized that Human Psychology rarely views gender in a vacuum—we naturally adjust our expectations based on the speaker's perceived age. They set out to replicate this inductive bias in a machine learning framework.
The "Raw" Methodology
The proposed system departs from the traditional pipeline of "Feature Extraction (MFCC) → Classifier" in two fundamental ways:
1. The Multi-Task Learning (MTL) Paradigm
Instead of just predicting "Male or Female," the network has two "heads."
- Primary Head: Predicts Gender.
- Auxiliary Head: Predicts the Age-Gender group (e.g., "Young Male" vs "Senior Female").
This forces the internal "hidden" layers of the AI to learn a representation that understands how age influences voice before it makes a final judgment on gender.
2. Learning from the Source
Most speech AI uses MFCCs—a "hand-crafted" summary of audio. This paper uses the Raw Waveform.
- Front-End: A 1D-Convolutional Neural Network (CNN) acts as a learnable filter bank, discovering acoustic features directly from the time-domain signal.
- Temporal Modeling: LSTMP (LSTM with Recurrent Projection) layers are used to track long-term patterns in the speech, outperforming the more common TDNN (Time Delay Neural Network) architectures.
Figure 1: The Multi-task learning DNN layout showing shared hidden layers and the dual-output head.
Experimental Proof
To test the model, the team combined data from NIST SRE (adults) and the OGI Kids corpus. The results were telling:
| Model | Weighted Accuracy (WA) | Unweighted Accuracy (UA) |
|---|---|---|
| Traditional GMM (MFCC) | 86.74% | 86.80% |
| Single-Task DNN (Raw Wave) | 87.89% | 87.77% |
| Multi-Task DNN (Raw Wave) | 90.70% | 90.65% |
The "Aha!" moment comes from the Ablation Study. When comparing the Single-Task model (Gender only) to the Multi-Task model (Gender + Age), the error rate dropped by 23.2%. This proves that the age information isn't just "extra data"—it's a critical guide for the network's understanding.
Table: Comparison shows the Multi-Task model generalizing significantly better across children and seniors.
Critical Insight & Future Outlook
This work highlights a shift in Speech AI: moving away from pre-defined mathematical features (like MFCC) toward end-to-end learning. More importantly, it shows that "Para-linguistic" traits (age, gender, emotion) are not independent.
Limitations: While gender identification improved, the authors noted that age classification did not necessarily see a boost from gender data. This suggests that while age is a precursor to understanding gender, the reverse might not be as numerically significant for the model.
Conclusion: For developers of Smart Call Centers or Human-Machine Interaction (HMI) systems, this paper provides a blueprint for building more inclusive AI that doesn't "misgender" a speaker simply because they are a child or an elderly person.
