Deciphering the Digital Pulse: A Multi-Modal Approach to Social Media Mood Identification
A Method to Identify the Current Mood of Social Media Users
The paper introduces a multi-modal mood identification system for social media users, categorizing emotions into Happy, Sad, Calm, and Angry. It utilizes a 1D Convolutional Neural Network (1D CNN) and a temporal weighted average of posts within 24 hours to achieve a classification accuracy of 85%.
TL;DR
This research presents a sophisticated framework to identify the "current mood" (Happy, Sad, Calm, Angry) of social media users. By combining 1D CNN-based text analysis, OCR for image-embedded text, and weighted emoticon scoring, the system captures the temporal nature of human emotion. With a weighted priority on recent posts, it achieves an impressive 85% accuracy.
Context: While sentiment analysis is a mature field, most models focus on static "opinions." This work shifts the focus toward dynamic "moods," positioning itself as a vital tool for mental health interventions and reactive recommendation systems.
Problem & Motivation: The Limitation of Binary Sentiments
Current Social Network Site (SNS) analysis often falls into the trap of oversimplification. Most public datasets only categorize content as Positive, Negative, or Neutral. However, human psychology is more complex; "Sadness" and "Anger" are both negative but lead to vastly different behaviors and recommendation needs.
Moreover, a user's mood isn't a permanent attribute. It fluctuates. Previous models often treated a post from 23 hours ago with the same weight as one from 5 minutes ago—silently ignoring the temporal decay of emotional states.
Methodology: Beyond Simple Keywords
The authors' approach is structured into three distinct sub-modules: data extraction (including OCR), scoring (Text + Emoticons), and temporal integration.
1. The Multi-Modal Input
The system doesn't just read status updates. It uses:
- Text Posts & Comments: Cleaned via custom pre-processing (HTML decoding, abbreviation handling).
- OCR Module: Using
pytesseractto extract text "baked" into shared images—a common way users express feelings today. - Emoticons: Weighted 60% in the final score (based on user surveys), recognizing that emojis often carry more emotional weight than the text itself.
2. 1D CNN Architecture
Moving beyond standard Neural Networks, the authors utilized a 1D Convolutional Neural Network. The rationale? CNNs are adept at identifying local patterns and n-gram sequences that signify specific moods, which simple dense layers often miss.

3. Temporal Weighting: The "Recency" Factor
One of the most insightful contributions is the Temporal Weighted Average. If a user makes four posts in a 24-hour window, the most recent post (Post 1) receives the highest weight (e.g., 4/10), while the oldest (Post 4) receives the lowest (1/10). This mimics the psychological reality that our current state is most heavily influenced by our most recent experiences.

Experiments & Results: Why 1D CNN?
The experimental results clearly justify the shift to 1D CNN. While a basic Neural Network suffered from significant overfitting (90% training vs. 64% testing), the 1D CNN maintained a robust testing accuracy of 83.88%, eventually peaking at 85% with hyperparameter tuning.
| Model | Training Accuracy | Testing Accuracy |
|---|---|---|
| Neural Network | 0.8999 | 0.6436 |
| 1D CNN | 0.9979 | 0.8388 |
The confusion matrix and F1 scores (averaging ~0.67) indicate that the model is balanced, though it occasionally struggles with "Sad" vs. "Angry" nuances—a common challenge in NLP due to overlapping lexicons in negative emotions.
Critical Analysis & Conclusion
Takeaway
The synergy of Text + Emoticons + Time is the secret sauce here. By weighting emoticons more heavily (0.6) and prioritizing recency, the model achieves a human-like "intuition" regarding the user's current state.
Limitations & Future Work
- Graphic Content: The current model "reads" images via OCR but doesn't "see" them. Future iterations should include Computer Vision (CV) to analyze colors and objects (e.g., dark colors for sadness).
- Data Scarcity: The dataset of 2,300 records is modest. Scaling this to a larger, more diverse demographic would likely improve the generalization of the "Calm" and "Angry" categories.
In conclusion, this methodology moves us closer to AI that truly understands the "vibe" of the digital world, offering a bridge between raw social media data and meaningful mental health support.
