Predictive Analytics for ADHD: Leveraging Machine Learning to Uncover Prevalence in Higher Education

Machine Learning Approach Applied to the Prevalence Analysis of ADHD Symptoms in Young Adults of Barranquilla, Colombia

2020-01-01
Alexandra Leon-Jacobus, Paola Patricia Ariza-Colpas, Ernesto Barcelo-Martínez, Marlon Alberto Piñeres-Melo, Roberto Cesar Morales-Ortega, David Alfredo Ovallos-Gazabon
Summary
Problem
Method
Results
Takeaways
Abstract

This study applies multiple Machine Learning (ML) classification algorithms to evaluate the prevalence of ADHD symptoms in 1,600 university students in Barranquilla, Colombia. The research identifies Bagging as the superior predictive model, achieving a state-of-the-art accuracy of 91.67% and a precision of 94.12% in distinguishing clinical symptoms using self-reporting instruments.

TL;DR

Attention Deficit Hyperactivity Disorder (ADHD) is frequently underdiagnosed in young adults, yet it significantly impacts academic retention and personal well-being. This research utilizes a massive dataset of 1,600 university students in Barranquilla, Colombia, applying advanced Machine Learning (ML) classifiers to identify prevalence. The study concludes that Bagging (Bootstrap Aggregation) is the most effective model, achieving a remarkable 91.67% accuracy in identifying ADHD symptoms.

Problem & Motivation: The "Hidden" ADHD in Adulthood

While ADHD is famously associated with childhood, as many as 70% of patients continue to exhibit symptoms into adulthood. However, the manifestation changes; the outward "hyperactivity" often transitions into internal disorganization, impulsivity, and affective lability.

The authors identify a critical gap: existing university screening methods often fail to capture the complexity of these symptoms, leading to poor academic performance or dropout. The motivation behind this study was to move beyond simple descriptive statistics and use Machine Learning to find patterns within 184 features of neuropsychological data that human observation might miss.

Methodology: The Power of Ensemble Learning

The research utilized a multi-stage approach, starting with a theoretical review and the administration of three specialized instruments:

  1. WURS (Wender Utah Rating Scale): Retrospective childhood symptom assessment.
  2. IES-ADHD: Evaluation of current inattention and hyperactivity.
  3. CIE-10 Test: Checklist based on WHO ICD-10 criteria.

Algorithm Benchmarking

The core of the methodology lies in the comparison of several high-performance ML algorithms, specifically focusing on Ensemble Methods. Ensemble methods combine multiple "weak" learners to create a "strong" learner.

  • Bagging: Reducer of variance through bootstrap sampling.
  • Boosting (MultiBoostAB, LogitBoost): Iteratively focuses on misclassified samples.
  • Decision Tree Variants (J48, DecisionStump): Structural mapping of attribute logic.

Overall Performance Metrics Table 1: Comparative performance of evaluated Machine Learning algorithms.

Experiments & Results: Bagging Takes the Lead

The experimentation phase used a dataset of 1,674 subjects, narrowed down to 1,600 after strict inclusion/exclusion filtering. Using the WEKA workbench, the authors evaluated the models based on Accuracy, Precision, Recall, and F-measure.

Key Findings:

  • Winner: Bagging achieved the highest accuracy (91.67%) and precision (94.12%).
  • The Power of Stability: Bagging's success confirms that for neuropsychological data—which can be "noisy" or unstable—ensemble methods that reduce variance are superior to single-tree models like J48.
  • Precision/Recall Balance: The F-measure (91.43%) indicates a very high reliability in balancing the identification of true positives without cluttering the results with false alarms.

Accuracy Comparison Chart Figure 1: Visualization of Accuracy across different ML architectures showing Bagging's superiority.

Critical Analysis & Conclusion

This work demonstrates that the transition from traditional cognitive assessment to Machine Learning-assisted diagnosis is not just possible, but highly accurate.

Takeaways for the Academic Community:

  • Scalability: This approach allows universities to screen thousands of students efficiently, identifying those who may need clinical intervention to succeed.
  • Theoretical vs. Empirical: The study validates that standard diagnostic criteria (DSM-V and ICD-10) can be effectively mapped into binary classification features for ML.

Limitations & Future Work:

The study relies on self-reported data, which can introduce bias. Future research should integrate biometric data (e.g., EEG or eye-tracking) with these ML classifiers to create a truly holistic "Neuro-AI" diagnostic tool.

By bridging the gap between psychiatry and computer science, this research paves the way for a future where student support systems are both data-driven and deeply personalized.

Find Similar Papers

Try Our Examples

  • Search for recent studies using ensemble machine learning techniques to predict ADHD prevalence in adult populations globally.
  • Which paper first established the Wender Utah Rating Scale (WURS) for adult ADHD assessment, and how do modern ML models improve its diagnostic accuracy?
  • Explore how Bagging and Boosting algorithms are being integrated into clinical decision support systems for other neurodevelopmental disorders like Autism or Dyslexia.
Contents
Predictive Analytics for ADHD: Leveraging Machine Learning to Uncover Prevalence in Higher Education
1. TL;DR
2. Problem & Motivation: The "Hidden" ADHD in Adulthood
3. Methodology: The Power of Ensemble Learning
3.1. Algorithm Benchmarking
4. Experiments & Results: Bagging Takes the Lead
4.1. Key Findings:
5. Critical Analysis & Conclusion
5.1. Takeaways for the Academic Community:
5.2. Limitations & Future Work: