Decoding the Soul via the Pen: Personality Classification Through Arabic Handwriting Analysis
Towards Personality Classification Through Arabic Handwriting Analysis
This paper introduces an automated framework for personality classification based on Arabic handwriting analysis using the Jung-Myers typology. By extracting language-specific features (such as Naskh and Ruqah styles) and applying machine learning algorithms like Decision Trees and Simple Logistic, the authors achieve a peak classification accuracy of 71.08% for the Judging/Perceiving dimension.
TL;DR
Can your handwriting reveal your deepest psychological traits? This study explores the intersection of Arabic Graphology and Machine Learning, leveraging the Jung-Myers typology (MBTI) to classify writers. By analyzing unique Arabic script features like Naskh and Ruqah styles, the researchers achieved up to 71.08% accuracy in predicting personality dimensions, proving that the way we loop our letters is a powerful biometric signal.
The Motivation: Moving Beyond Manual Graphology
Graphology—the study of handwriting to determine personality—has existed for centuries but has long been criticized for its lack of scientific rigor and reliance on human intuition. In the digital age, the need for automated, objective analysis is surging in fields like Human Resources, Forensics, and Criminal Justice.
While significant progress has been made in Latin script recognition, Arabic handwriting presents a unique set of challenges:
- High Character Similarity: Many letters differ only by the placement of dots.
- Cursive Nature: Arabic is inherently cursive, making segmentation and feature extraction difficult.
- Stylistic Diversity: The coexistence of classical styles like Naskh (formal) and Ruqah (shorthand/abbreviated) introduces a layer of complexity absent in English.
Figure 1: The standardized form used for data collection, featuring both normal and fast writing tasks.
Methodology: Bridging Psychology and Feature Engineering
The authors adopted the Jung-Myers Typology, which categorizes personalities into four dichotomies:
- E vs. I: Extraversion vs. Introversion (Main Attitude)
- S vs. N: Sensation vs. Intuition (Perceptive Function)
- T vs. F: Thinking vs. Feeling (Judging Function)
- J vs. P: Judging vs. Perceiving (Attitude toward the outside world)
Feature Extraction
The core of the methodology lies in the translation of visual handwriting artifacts into mathematical features. The authors identified eight key features:
- Geometric Features: Baseline slope, letter slant, and word spacing.
- The "Arabic" Insight: The inclusion of Writing Style. The choice between Naskh (meticulous) and Ruqah (efficient/abbreviated) is hypothesized to reflect personality-driven cognitive biases (e.g., detail-oriented vs. efficiency-oriented).
Machine Learning Pipeline
The researchers compared four distinct classification algorithms:
- Simple Logistic (SL)
- Decision Tree (J48)
- K-Nearest Neighbors (IBK)
- Random Forest (RF)
Experimental Results: The Power of Script
The study conducted two main experiments. Experiment 1 included both handwriting features and demographic metadata (age, gender, etc.), while Experiment 2 relied solely on handwriting.
Surprisingly, the results across both experiments were nearly identical, suggesting that personality traits are encoded directly in the script, independent of the writer's nationality or age.
| Dimension | Best Algorithm | Peak Accuracy |
|---|---|---|
| E / I | J48 (Decision Tree) | 68.67% |
| J / P | J48 (Decision Tree) | 71.08% |
| T / F | IBK (KNN) | 61.44% |
| S / N | Simple Logistic | 56.62% |
Figure 2: Performance comparison across different machine learning algorithms. The Decision Tree (J48) consistently outperformed others in the Attitude dimensions.
Critical Analysis & Conclusion
Takeaways
The most significant contribution of this work is the validation of Arabic-specific features like the Naskh/Ruqah style in psychological profiling. The fact that the Judging vs. Perceiving (J/P) dimension saw the highest accuracy (71.08%) aligns with graphological theory: "Judging" individuals tend to be more structured and consistent in their script, while "Perceiving" individuals may show higher variance and speed-related distortions.
Limitations & Future Outlook
The primary bottleneck cited is the limited dataset size (83 participants). For machine learning, specifically in high-dimensional tasks like personality classification, a few hundred samples are insufficient to capture the full spectrum of the 16 MBTI types.
Future Work should look toward Deep Learning (CNNs and LSTMs). Rather than manually defining features like "slant" or "margin," deep neural networks could learn latent features directly from raw image pixels, potentially uncovering patterns that are invisible to the human eye.
In conclusion, while your handwriting might not yet serve as a definitive psychological profile, this research brings us one step closer to an era where the stroke of a pen is as revealing as a DNA test.
