CRF vs. HELM: Advancing Stuttering Detection in Children's Speech via Data Augmentation

Detecting Stuttering Events in Transcripts of Children’s Speech

2017-01-01
Sadeen Alharbi, Madina Hasan, Anthony J. H. Simons, Shelagh Brumfitt, Phil D. Green
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates the automated detection of six types of stuttering events in transcripts of children's speech using Hidden Event Language Models (HELM) and Conditional Random Fields (CRF). The study demonstrates that CRF achieves superior performance (94% accuracy) and highlights the critical role of data augmentation in addressing the scarcity of disfluent speech data.

Executive Summary

TL;DR: This research addresses the challenge of identifying speech disorders in children by automating stuttering event detection in text transcripts. By comparing Conditional Random Fields (CRF) and Hidden Event Language Models (HELM), and introducing a massive synthetic dataset for augmentation, the authors achieved a high-accuracy (94%) classification system capable of identifying even rare stuttering types like prolongations.

Background: This work moves beyond traditional, rigid rule-based systems into the territory of probabilistic sequence labeling. It sits at the intersection of Clinical Speech-Language Pathology and Natural Language Processing (NLP), significantly expanding the UCLASS (University College London Archive of Stuttered Speech) corpus.

Problem & Motivation

Early intervention is critical for children who stutter due to higher neural plasticity at a young age. However, clinical assessment involves tedious manual counting of disfluencies, which is:

  1. Subjective: Dependent on the clinician’s experience.
  2. Time-Intensive: Transcribing and categorizing every word takes significant effort.

Previous automated attempts used Rule-Based (RB) algorithms. While precise for specific cases, they are deterministic and fail when encountering speech patterns not explicitly predefined by experts. Furthermore, the field suffers from a chronic lack of training data, especially for specific disfluency types like "Prolongations" or "Part-word repetitions."

Methodology: The Shift to Sequence Labeling

The authors treat stuttering as a "hidden event" occurring within or between words. They compared two primary architectures:

  1. HELM (Hidden Event Language Model): A generative approach using quasi-HMM techniques to estimate the probability of a stuttering event following a word given its context.
  2. CRFs (Conditional Random Fields): A discriminative model that optimizes the posterior probability of a label sequence based on local features (n-grams and "post-words").

Data Augmentation Strategy

To solve the data scarcity problem, the team used the SRILM toolkit to generate "nonsense" but statistically plausible sentences weighted by the stuttering patterns found in the training set. This expanded the available data from ~13,000 words to over 400,000 words.

Model Approach and Data Distribution Table: Distribution of stuttering types. Note the initial scarcity of Phrase Repetitions (PH) and Prolongations (P).

Experiments & Results

The study found that local context is king. For CRF, the optimal configuration used 2-grams combined with 2-post-words (looking two words ahead).

Key Findings:

  • CRF Dominance: CRF consistently outperformed HELM, likely due to its ability to handle overlapping features and its discriminative nature.
  • The Augmentation Breakthrough: Before augmentation, both models had a 0% Recall for Prolongations (P) and Part-word repetitions (PW). After adding synthetic data, CRF was able to detect these with significant precision (see Table below).

Baseline vs Augmented Results Table: The post-augmentation performance shows a dramatic rise in F1-scores for rare classes like P and PW.

Error Analysis (Confusion Matrix)

While the system reached 94% accuracy, it still struggled with Phrase Repetitions (PH). The authors noted that their augmentation method was word-based rather than phrase-based, meaning the synthetic data didn't provide enough examples of multi-word stuttering patterns.

Critical Analysis & Conclusion

Takeaway: This paper proves that even "nonsense" synthetic data can be incredibly valuable in medical NLP tasks if it preserves the local statistical structure of the disorder. It successfully shifts the burden of diagnosis from rigid rules to adaptable machine learning models.

Limitations:

  • Phrase Patterns: The current augmentation lacks phrase-level logic.
  • ASR Dependency: The study assumes clean transcripts. In a real-world clinical setting, an ASR system would first need to transcribe the child's speech, potentially introducing its own errors.

Future Work: The next frontier involves a joint model that handles both the audio signal (for sound-level dysfluencies) and the transcript (for lexical-level dysfluencies) simultaneously to create a truly end-to-end diagnostic tool.

Find Similar Papers

Try Our Examples

  • Find recent papers that apply deep learning sequence models like BiLSTM-CRF or Transformers to the task of stuttering event detection in speech transcripts.
  • What are the primary methodologies for "Vocal Tract Length Perturbation" and other audio-level data augmentation techniques mentioned as precursors to this text-based study?
  • Search for studies that integrate Automatic Speech Recognition (ASR) confidence scores with stuttering detection models to handle errors in children's speech transcription.
Contents
CRF vs. HELM: Advancing Stuttering Detection in Children's Speech via Data Augmentation
1. Executive Summary
2. Problem & Motivation
3. Methodology: The Shift to Sequence Labeling
3.1. Data Augmentation Strategy
4. Experiments & Results
4.1. Key Findings:
4.2. Error Analysis (Confusion Matrix)
5. Critical Analysis & Conclusion