Takemaru-kun: Bridging the Generative Gap Between Adult and Child Speech in Public Spaces

Public speech-oriented guidance system with adult and child discrimination capability

2004-09-28
Ryuichi Nisimura, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces an age-group discrimination system for the "Takemaru-kun" public speech guidance agent, utilizing a Support Vector Machine (SVM) to distinguish between adult and child users. By combining acoustic likelihoods with linguistic scores from parallel decoders, the system achieves a 92.4% discrimination accuracy, enabling age-appropriate dialogue responses.

TL;DR

Deploying speech interfaces in the "wild" (like community centers) reveals a harsh reality: children often dominate usage, yet adult-trained models fail them. This paper presents a robust discrimination system for the Takemaru-kun guidance agent that uses parallel decoding and SVMs to distinguish between adults and children with 92.4% accuracy, paving the way for adaptive, age-appropriate AI interactions.

Background: The "Child Problem" in Public Speech Systems

When the Takemaru-kun system was installed in Ikoma-city, the researchers discovered that 68.2% of users were children. However, the system's ability to provide appropriate responses plummeted from 73.4% for adults to a meager 37.4% for children. This gap stems from two factors:

  1. Acoustic Mismatch: Children have higher fundamental frequencies and different vocal tract lengths.
  2. Linguistic Divergence: As shown in the paper's analysis, children ask about the agent's "personal life" or use greetings, while adults seek specific "Guidance" information.

Methodology: Parallel Decoders and Multi-Feature Fusion

The core innovation is not just looking at how someone sounds (acoustics), but what they say (linguistics).

1. The Dual-Decoder Architecture

The system runs two speech recognition engines in parallel:

  • Adult Model: Trained on JNAS male/female data and adapted to natural adult utterances via MAP/MLLR.
  • Child Model: Specifically adapted using 17,000+ utterances from actual child users.

2. Feature Extraction

Instead of raw audio, the system extracts high-level statistical features from the decoders:

  • Average Acoustic Score (AP): The log-likelihood of the acoustic model normalized by frames.
  • Language Score (LP): The log-likelihood of the trigram language model normalized by word count.

Model Architecture and Takemaru-kun Interface Fig 1: The Takemaru-kun physical setup and the animation agent interface.

Overcoming the "Female-Child" Overlap

A classic failure mode in age discrimination is the acoustic similarity between adult female voices and children. The authors show that using MAP (Maximum A Posteriori) Adaptation effectively shifts the acoustic models to create a clear separation in the likelihood space.

Acoustic Distribution Comparison Fig 2: Distributions of AP scores showing how MAP adaptation successfully separates adult female signals from child signals.

Experimental Results

The researchers compared their SVM-based approach against a standard Gaussian Mixture Model (GMM) baseline.

MethodDiscrimination Rate
GMM Baseline86.4%
SVM (Acoustic Only)91.6%
SVM (Acoustic + Linguistic)92.4%

The inclusion of Linguistic Properties (LP) proved vital. Because children tend to use a specific subset of "out-of-task" language or different greeting patterns, the language model likelihood acts as a powerful secondary classifier.

Word Accuracy Table Fig 3: Word accuracy improvements after adaptation highlight the necessity of age-specific models.

Critical Insight & Future Outlook

This work demonstrates that for real-world "In-the-Wild" AI, context is king. By treating the speech recognition process as a feature generator for an SVM classifier, the authors created a system that doesn't just recognize words—it recognizes the type of human behind the words.

Future Work: The authors suggest exploring even more granular machine learning algorithms and expanding the system to handle varied dialogue strategies (e.g., using simpler language when a child is detected). As we move toward more ubiquitous voice assistants, this "age-aware" Inductive Bias will be crucial for inclusive UX design.

Conclusion

The Takemaru-kun system provides a roadmap for public-facing speech interfaces. By leveraging the synthesis of acoustic and linguistic data, we can move past "one-size-fits-all" models toward systems that truly understand their audience.

Find Similar Papers

Try Our Examples

  • Find recent papers on age-group classification in speech recognition that utilize deep learning or transformer-based embeddings instead of log-likelihood scores.
  • Searching for the foundational research on Phonetic Tied Mixture (PTM) models and how they have evolved into modern neural acoustic modeling.
  • What are the current SOTA methods for multi-modal speaker diarization in public kiosks or robotic guidance systems?
Contents
Takemaru-kun: Bridging the Generative Gap Between Adult and Child Speech in Public Spaces
1. TL;DR
2. Background: The "Child Problem" in Public Speech Systems
3. Methodology: Parallel Decoders and Multi-Feature Fusion
3.1. 1. The Dual-Decoder Architecture
3.2. 2. Feature Extraction
4. Overcoming the "Female-Child" Overlap
5. Experimental Results
6. Critical Insight & Future Outlook
7. Conclusion