The Emotional Keyboard: Deciphering Mood via Keystrokes and Motion
Emotion recognition using mobile phones
This paper presents an intelligent emotion detection system for Android mobile phones, implemented as a custom soft keyboard. Using machine learning (specifically the J48 algorithm), the system predicts four emotional states (Neutral, Angry, Happy, Sad) by analyzing accelerometer data and typing behaviors such as speed and backspace usage, achieving over 90% accuracy.
TL;DR
Researchers have developed a smart Android keyboard that knows how you feel just by how you type, not what you type. By analyzing typing speed, errors, and phone movement using the J48 algorithm, the system achieves a staggering 90.2% average accuracy in distinguishing between Happy, Sad, Angry, and Neutral states—all without reading your private messages.
Background: Beyond Survey Fatigue
Capturing human emotion digitally has long been a "holy grail" for UX researchers and healthcare providers. Historically, we've relied on two flawed methods:
- Self-Reporting: Asking users "How do you feel?" (interruptive and biased).
- NLP (Natural Language Processing): Analyzing text for sentiment (fails at sarcasm, slang, and is privacy-invasive).
This paper suggests a third way: Implicit Behavioral Sensing. The way your hand shakes when you are angry or the long pauses you take when sad are biometric signatures that are hard to fake and easy for sensors to catch.
The Methodology: The Mechanics of a "Sad" Backspace
The authors replaced the standard Android Open Source Project (AOSP) keyboard with a custom version that logs data in 5-second segments.
Feature Engineering
The model focuses on three primary non-linguistic features:
- TimeBP (Time Between Presses): The millisecond delay between keystrokes.
- Backspaces: A proxy for cognitive load or frustration-induced errors.
- Accelerometer Data: Captured in three dimensions to detect hand shakiness or grip intensity.
Algorithm Showdown
The team evaluated five heavyweights from the machine learning world: Naïve Bayes, J48 Decision Trees, IBK (Nearest Neighbor), Linear Regression, and SVM. Surprisingly, SVM failed miserably (ROC Area ~0.60), while J48 excelled.
Fig 1: The data capture and prediction cycle within the Android environment.
Experimental Insights: Why J48 Won
The failure of SVM suggests that human emotions don't sit on opposite sides of a simple straight line. Instead, the data is "clumpy." The J48 decision tree (shown below) reveals that "Angry" can be detected almost purely through extreme accelerometer movement, while "Happy" vs. "Neutral" requires a more nuanced check of typing speed (TimeBP).
Fig 2: The logic behind the machine's "feelings"—note how TimeBP is a critical node for Happy/Neutral differentiation.
Key Performance Metrics (J48)
| Class | Precision | Recall | F-Measure |
|---|---|---|---|
| Angry | 1.000 | 1.000 | 1.000 |
| Neutral | 0.905 | 0.872 | 0.888 |
| Happy | 0.821 | 0.821 | 0.821 |
| Sad | 0.904 | 0.979 | 0.940 |
| Average | 0.902 | 0.902 | 0.902 |
Critical Analysis & Future Outlook
The beauty of this system is its application independence. Because the sensor data is collected at the keyboard level, it works whether you are writing an email, tweeting, or searching for a recipe.
Limitations
- Small Sample Size: The model was trained on 307 feature vectors from 3 users. While the cross-validation is promising, real-world diversity in typing styles might lower accuracy.
- Hardware Variance: Different phones have different accelerometer sensitivities.
The "Big Picture"
The authors also implemented a web-based "GeoChart" to track regional happiness or sadness. Imagine a world where public health officials can see a city's stress levels rising in real-time during a crisis—without ever reading a single private text. This keyboard is a step toward that "Emotionally Aware" digital ecosystem.
Summary Takeaway
By shifting from what we say to how we interact with our devices, we unlock a privacy-respecting and highly accurate window into the human psyche. The J48 classifier proves that simple sensor data, when structured correctly, can outperform complex linguistic analysis.
