Identifying At-Risk Students: Leveraging Decision Trees for Early Dropout Prediction in Blended Learning

Mining Educational Data to Predict Academic Dropouts: a Case Study in Blended Learning Course

2018-10-01
Otgontsetseg Sukhbaatar, Kohichi Ogata, Tsuyoshi Usagawa
Summary
Problem
Method
Results
Takeaways
Abstract

This study proposes a dropout prediction scheme using Decision Tree analysis within a Blended Learning environment. By extracting four key quiz-related features from Learning Management System (LMS) logs, the researchers developed a classifier capable of identifying dropout-prone students by the middle of the semester with an overall accuracy of up to 89%.

TL;DR

This research addresses the critical challenge of student retention in higher education by mining LMS data from a sophomore engineering course. By focusing on quiz-related behavioral features—such as persistence and engagement—the authors developed a Decision Tree-based warning system that identifies nearly 80% of dropout-prone students by mid-semester, allowing for timely pedagogical intervention.

Background & Positioning

In the landscape of Educational Data Mining (EDM), this work sits at the intersection of behavioral analysis and predictive modeling. While many SOTA models focus on massive datasets from MOOCs, this study targets the "Blended Learning" niche—where traditional face-to-face lectures are augmented by online components. Its core value lies in proving that historical data from previous iterations of a course can effectively predict outcomes for current students, provided the course structure remains stable.

Problem & Motivation

Predicting academic failure is inherently harder than predicting commercial web behavior. In a blended learning environment, the "digital footprint" of a student is often sparse. The authors identified a specific gap: instructional timing. Most models analyze data post-facto, but for a prediction to be useful, it must occur early enough (mid-semester) for an instructor to intervene. The researchers hypothesized that "persistence" in online quizzes—the main online activity—would serve as a proxy for a student's commitment to the course.

Methodology: The Power of Quiz-Related Features

The authors moved away from simple "total clicks" and instead extracted four high-signal features related to quiz behavior:

  1. Quiz Attempt Frequency: The rhythm of study.
  2. Quiz Attempt Duration: Identifying "guess tries" versus "problem-solving" effort.
  3. Total Activity Count: Overall engagement level.
  4. Number of Quiz Attempts: An indicator of persistence.

Using a Decision Tree (Gini Index), they established specific thresholds (e.g., frequency > 11.16 days or activities < 104) to flag students.

Model Feature Criteria Figure 1: Decision Tree Classifiers derived from 2012 historical data used for predictive rules.

Experimental Results

The "True Positive" rate for dropouts was impressive. In the 2013 cohort, 100% of dropouts were caught, and in 2016, 58% were identified. The overall accuracy remained high at 87-89%.

However, the study revealed a significant limitation: predicting "Failure" (completing but not passing) is significantly harder than predicting "Dropout".

Performance Comparison Table Figure 2: Accuracy metrics including Sensitivity, Precision, and F-score across 2013 and 2016 test sets.

The data shows that failing students often put in as much (or more) effort in the LMS as successful ones; they simply struggle with the content. This suggests that behavioral logs alone cannot solve the "failure" problem; content-mastery metrics are required.

Critical Insight & Conclusion

Takeaways

  • The Stability Advantage: In university settings where courses are fundamental (like Engineering Math), last year's data is a goldmine for this year's predictions.
  • Activity vs. Mastery: Log data is excellent at detecting "disengagement" (dropout) but poor at detecting "struggle" (failure).

Limitations & Future Work

The 8th week may still be too late for some. Future iterations should aim for weekly tracking to catch the "rapid disengagers." Furthermore, integrating the scores of the quizzes rather than just the frequency of attempts might help bridge the gap in identifying failing students.

Ultimately, this work provides a robust, low-complexity framework that any university instructor using an LMS (like Moodle or Canvas) could implement to save students from falling through the cracks.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize Deep Learning (e.g., RNNs or LSTMs) to improve the prediction of failing students in blended learning models where basic behavioral metrics are insufficient.
  • Which original papers first established the "Gini Index" as the primary metric for Decision Tree classification in Educational Data Mining, and how has its use evolved?
  • Explore research that applies dropout prediction frameworks from traditional university LMS data to the specific time-series dynamics of Massive Open Online Courses (MOOCs).
Contents
Identifying At-Risk Students: Leveraging Decision Trees for Early Dropout Prediction in Blended Learning
1. TL;DR
2. Background & Positioning
3. Problem & Motivation
4. Methodology: The Power of Quiz-Related Features
5. Experimental Results
6. Critical Insight & Conclusion
6.1. Takeaways
6.2. Limitations & Future Work