Predictive Analytics for ADHD: Leveraging Machine Learning to Uncover Prevalence in Higher Education
Machine Learning Approach Applied to the Prevalence Analysis of ADHD Symptoms in Young Adults of Barranquilla, Colombia
This study applies multiple Machine Learning (ML) classification algorithms to evaluate the prevalence of ADHD symptoms in 1,600 university students in Barranquilla, Colombia. The research identifies Bagging as the superior predictive model, achieving a state-of-the-art accuracy of 91.67% and a precision of 94.12% in distinguishing clinical symptoms using self-reporting instruments.
TL;DR
Attention Deficit Hyperactivity Disorder (ADHD) is frequently underdiagnosed in young adults, yet it significantly impacts academic retention and personal well-being. This research utilizes a massive dataset of 1,600 university students in Barranquilla, Colombia, applying advanced Machine Learning (ML) classifiers to identify prevalence. The study concludes that Bagging (Bootstrap Aggregation) is the most effective model, achieving a remarkable 91.67% accuracy in identifying ADHD symptoms.
Problem & Motivation: The "Hidden" ADHD in Adulthood
While ADHD is famously associated with childhood, as many as 70% of patients continue to exhibit symptoms into adulthood. However, the manifestation changes; the outward "hyperactivity" often transitions into internal disorganization, impulsivity, and affective lability.
The authors identify a critical gap: existing university screening methods often fail to capture the complexity of these symptoms, leading to poor academic performance or dropout. The motivation behind this study was to move beyond simple descriptive statistics and use Machine Learning to find patterns within 184 features of neuropsychological data that human observation might miss.
Methodology: The Power of Ensemble Learning
The research utilized a multi-stage approach, starting with a theoretical review and the administration of three specialized instruments:
- WURS (Wender Utah Rating Scale): Retrospective childhood symptom assessment.
- IES-ADHD: Evaluation of current inattention and hyperactivity.
- CIE-10 Test: Checklist based on WHO ICD-10 criteria.
Algorithm Benchmarking
The core of the methodology lies in the comparison of several high-performance ML algorithms, specifically focusing on Ensemble Methods. Ensemble methods combine multiple "weak" learners to create a "strong" learner.
- Bagging: Reducer of variance through bootstrap sampling.
- Boosting (MultiBoostAB, LogitBoost): Iteratively focuses on misclassified samples.
- Decision Tree Variants (J48, DecisionStump): Structural mapping of attribute logic.
Table 1: Comparative performance of evaluated Machine Learning algorithms.
Experiments & Results: Bagging Takes the Lead
The experimentation phase used a dataset of 1,674 subjects, narrowed down to 1,600 after strict inclusion/exclusion filtering. Using the WEKA workbench, the authors evaluated the models based on Accuracy, Precision, Recall, and F-measure.
Key Findings:
- Winner: Bagging achieved the highest accuracy (91.67%) and precision (94.12%).
- The Power of Stability: Bagging's success confirms that for neuropsychological data—which can be "noisy" or unstable—ensemble methods that reduce variance are superior to single-tree models like J48.
- Precision/Recall Balance: The F-measure (91.43%) indicates a very high reliability in balancing the identification of true positives without cluttering the results with false alarms.
Figure 1: Visualization of Accuracy across different ML architectures showing Bagging's superiority.
Critical Analysis & Conclusion
This work demonstrates that the transition from traditional cognitive assessment to Machine Learning-assisted diagnosis is not just possible, but highly accurate.
Takeaways for the Academic Community:
- Scalability: This approach allows universities to screen thousands of students efficiently, identifying those who may need clinical intervention to succeed.
- Theoretical vs. Empirical: The study validates that standard diagnostic criteria (DSM-V and ICD-10) can be effectively mapped into binary classification features for ML.
Limitations & Future Work:
The study relies on self-reported data, which can introduce bias. Future research should integrate biometric data (e.g., EEG or eye-tracking) with these ML classifiers to create a truly holistic "Neuro-AI" diagnostic tool.
By bridging the gap between psychiatry and computer science, this research paves the way for a future where student support systems are both data-driven and deeply personalized.
