Analyzing Student Mental Health: A Deep Learning Approach for the Moroccan Context
Assessment of Lifestyle and Mental Health: Case Study of the FST Beni Mellal
The paper presents a multi-class sentiment analysis framework to assess student lifestyle and mental health at FST Beni Mellal, Morocco. It develops a custom deep learning pipeline, culminating in a CNN-based model (utilizing GloVe embeddings) that classifies social media text into 13 emotional categories, achieving a SOTA-level accuracy of 74% for this specific context.
TL;DR
Researchers at Sultan Moulay Slimane University have developed a customized text classification pipeline to monitor student mental health by analyzing social media posts. By evolving from traditional Machine Learning to a refined CNN-based architecture using GloVe embeddings, they achieved a 74% accuracy in classifying complex emotional states within the "Darija" dialect context, significantly outperforming standard recurrent models.
Context & Motivation: The "Darija" Challenge
Monitoring mental health through digital footprints (Web Scraping) offers a non-intrusive, real-time alternative to surveys. However, the linguistic landscape at the Faculty of Sciences and Technologies (FST) Beni Mellal is unique. Students communicate in Darija, a Moroccan dialect that lacks rigid spelling rules and often utilizes a mix of Arabic and French characters.
Existing SOTA models trained on Formal Arabic or French often fail to capture the nuances of this "Internet-Darija." The authors' primary motivation was to move beyond simple sentiment (Positive/Negative) to a 13-class emotional spectrum (Anger, Joy, Sadness, etc.) to better reflect psychological well-being.
Methodology: From Baselines to Deep Learning
The research followed a progressive path, testing three distinct architectural approaches:
1. The Conventional Baseline (ML + Grid Search)
The authors first established a baseline using SVM, KNN, and Naïve Bayes. A critical insight here was the use of Grid Search to optimize the SVM, which saw a massive jump from 50.7% to 72.9% accuracy after adjusting the regularization parameter (C) and the RBF kernel.
2. The Deep Learning Evolution
To surpass the 72.9% threshold, three DL models were constructed:
- Model 1 (LSTM-RNN): Used Long Short-Term Memory cells with an embedding layer. Surprisingly, it underperformed (49% accuracy), likely due to the limited dataset size not sufficing for the complex temporal dependencies in short tweets.
- Model 2 (Transfer Learning): Leveraged a model pre-trained on the Reuters news corpus. By stripping the output layer and retraining it for 13 classes, accuracy rose to 72%.
- Model 3 (CNN + GloVe): The "Gold Standard" of the study. By using 1D-Convolutions and GloVe (Global Vectors for Word Representation), the model effectively captured local keyword patterns essential for emotional detection in short-form text.
Figure 1: The overall structural framework from Data Input to Prediction.
Experimental Results & Insights
The Comparison between the models highlights a critical trend: CNNs are highly effective for short-text classification where specific "trigger words" carry significant emotional weight.
| Model | Precision | Recall | F-Measure | Accuracy |
|---|---|---|---|---|
| First Model (LSTM) | 24% | 49% | 32% | 49% |
| Second Model (Transfer) | 70% | 72% | 70% | 72% |
| Third Model (CNN) | 72% | 74% | 71% | 74% |
One of the most revealing findings was the failure of the LSTM model. In many NLP tasks, LSTMs are preferred for their sequence handling, but in the context of erratic dialectal social media posts, the CNN's ability to extract n-gram-like features (via Kernels) proved more robust.
Table 1: Performance metrics showing the superiority of the CNN architecture.
Critical Analysis & Future Outlook
Strengths:
- Contextual Relevance: Unlike many general studies, this work focuses on a hyper-local linguistic context (Moroccan Darija).
- Architecture Hybridization: Combining CNNs with pre-trained GloVe vectors successfully mitigated the "small dataset" problem.
Limitations:
- Dictionary Constraints: The authors noted that "Darija" currently lacks a specialized NLP system for lemmatization and stemming, which likely capped the accuracy at 74%.
- Data Volume: Only a hundredth of the scraped 3,500 posts were manually verified due to team size constraints.
The Path Forward: The authors propose the development of an Autonomous NLP system for Darija using Autoencoders and Reinforcement Learning. Such a system would allow for sentence-level semantic understanding rather than just keyword-level sentiment analysis, potentially pushing accuracy into the 80-90% range seen in English-centric SOTA models.
Conclusion
This case study at FST Beni Mellal serves as a blueprint for localized mental health monitoring. It proves that even with linguistic complexities like Darija, a well-tuned CNN-GloVe pipeline can provide actionable insights into the emotional state of a community.
