Automated Linguistic Profiling: Distinguishing ASD, Down Syndrome, and Typical Development in Children
Analysis of Dialogues of Typically Developing Children, Children with Down Syndrome and ASD Using Machine Learning Methods
This paper presents a machine learning-based approach to differentiate between the speech of Typically Developing (TD) children, children with Autism Spectrum Disorder (ASD), and children with Down Syndrome (DS). Using automatic graphemic and morphological analysis of transcribed dialogues, the researchers extracted 25 linguistic features and employed tree-based ensembles (Gradient Boosting, Random Forest, AdaBoost) to achieve a classification accuracy of 83%.
Executive Summary
TL;DR: Researchers have developed an automated pipeline that uses machine learning to "read" the differences in how children with ASD, Down Syndrome (DS), and typical development (TD) speak. By analyzing morphological signatures—like the frequency of verbs vs. nouns or the length of sentences—the system can classify a child's developmental profile with 83% accuracy.
Field Positioning: This study moves beyond traditional manual linguistics into the realm of Computational Psycholinguistics, providing evidence that automated NLP tools designed for standard Russian are effective for analyzing atypical speech patterns.
The "Why": Why Manual Analysis Falls Short
Understanding speech development is crucial for early intervention. However, the current "Gold Standard" involves linguists manually transcribing and tagging every noun, verb, and pause. This is slow and prone to subjective error. Furthermore, existing research often looks at ASD in isolation; this study seeks a broader "differential diagnosis" by comparing TD, ASD, and DS children simultaneously within the same experimental framework.
Methodology: From Raw Text to Feature Vectors
The researchers utilized a dataset of 62 Russian-speaking boys (8-11 years old). The method follows a three-stage pipeline:
- Graphematic Analysis: Extracting structural units like tokens (words), sentences, and pauses.
- Morphological Tagging: Using automated tools to identify 16 different parts of speech (ADJF, NOUN, VERB, INTJ, etc.).
- Statistical Pruning: Using the Kruskal-Wallis H Test to identify which features actually matter (reducing 25 features down to the 12 most significant ones).

The core insight here is that morphological richness serves as a proxy for cognitive processing strategies. For example, children with DS showed higher frequencies of "interjections" and "non-existent words," reflecting specific phonetic and lexical challenges.

Results & Machine Learning Performance
The researchers tested several ensemble models. While individual models like Gradient Boosting and AdaBoost performed well, a Voting Classifier (which combines the predictions of multiple models) proved superior.
- Typical Development (TD): Highest accuracy (95% recall). Their speech is characterized by more tokens and significantly higher usage of conjunctions and adverbs.
- ASD: 86% recall. Their attributes often fall "in the middle" between TD and DS.
- Down Syndrome (DS): 57% recall. This was the hardest group to classify, often being confused with ASD due to overlapping simplified sentence structures.

Deep Insight: The Spectrum of Complexity
One of the most profound takeaways is the Gradient of Complexity. The Kruskal-Wallis tests revealed that features like "relative frequency of conjunctions" and "tokens per sentence" follow a strict descending order: TD > ASD > DS. This suggests that automated linguistic analysis isn't just detecting "errors"; it's mapping the structural complexity of a child's internal language model.
Critical Analysis & Future Outlook
Limitations: The sample size is relatively small (62 children), and the model struggles to differentiate between DS and ASD. This is likely because both conditions can result in shorter, simplified sentences, making "breadth of vocabulary" a noisy feature.
The Future: The authors suggest adding Acoustic Analysis (pitch, rhythm, tone) and more granular grammatical features (verb tenses, noun animation) to the mix. Combining "how they say it" with "what they say" could provide the necessary signal to distinguish ASD from DS with even higher precision.
Conclusion
This paper serves as a proof-of-concept for automated screening tools. By leveraging standard NLP morphological analyzers, we can build scalable, objective systems to support clinicians in tracking child development and identifying speech-language pathologies earlier than ever before.
