Deep Learning vs. Traditional ML: Shielding Gender Secrets in Russian Texts
Deep Learning neural nets versus traditional machine learning in gender identification of authors of RusProfiling texts
The paper presents a comparative study of gender identification in Russian texts using the RusProfiling and RusPersonality corpora, pitting conventional machine learning (SVM, Gradient Boosting) against Deep Learning architectures. The authors introduce "Model 1," a multi-layer Convolutional Neural Network (CNN) that achieves a state-of-the-art F1-score of 88% on the gender profiling task.
TL;DR
In the evolving landscape of stylometry, identifying the gender of an author behind a screen is a classic yet complex task. This paper investigates whether modern Deep Learning can surpass the rigorous baselines set by traditional Machine Learning on Russian corpora. The verdict: Model 1, a deep CNN, dominates with an 88% F1-score, setting a new benchmark for the RusProfiling task and proving that hidden stylistic signatures can be extracted more effectively through deep neural hierarchies than through manual feature engineering.
Problem & Motivation
Russian is a morphologically rich language where gender often manifests clearly through verb endings and adjective suffixes. While this sounds like a "solved" classification problem, the difficulty spikes when:
- Authors intentionally deceive: Writing as the opposite sex to mask identity.
- Short, informal texts: Social media (Twitter, FB) lacks the formal structures that traditional SVMs rely on.
- Domain Adaptation: Models trained on letters often fail when tested on blog posts or reviews.
The authors' intuition was that while traditional ML (SVM/Gradient Boosting) captures "surface" statistics, Deep Learning—specifically CNNs and LSTMs—could capture the rhythm and latent structural patterns of gendered writing that persist even when surface pointers are altered.
Methodology: The Core Architectures
The researchers tested several "data-driven" modeling pipelines. The most significant shift was moving from character n-grams to sophisticated neural topologies.
The CNN Powerhouse (Model 1)
Unlike many NLP tasks that default to Recurrent Neural Networks (RNNs), this paper highlights a deep CNN approach. The architecture utilizes four sequential convolutional layers with ReLU activations followed by Global Max Pooling. This design is specialized for detecting local patterns (like specific phrase turns or punctuation habits) regardless of their position in the text.
The LSTM Contenders (Models 2 & 3)
Stacked Bidirectional LSTMs were deployed to capture long-range dependencies. These models process text both forward and backward, attempting to understand the context surrounding every word or morphological tag.
Note: The study leveraged a diverse set of features, from Word2vec to Morphological vectors (GRM-1/2).
Experiments & SOTA Results
The comparison results were decisive. Throughout various splits of the RusProfiling (Tw, FB, LJ) and RusPersonality corpora, Deep Learning maintained a significant lead.
| Model | Feature Set | F1-Score | Delta vs. Baseline |
|---|---|---|---|
| Model 1 (CNN) | Seq. feat | 0.88 | +38% |
| Gradient Boosting | TF-IDF | 0.79 | +29% |
| SVM | Char N-grams | 0.74 | +15% |
Table 6: Performance on balanced training sets shows Model 1's clear superiority.
Key Insights from Experiments:
- More Data, Better Performance: As seen in the SVM learning curve (Figure 1), traditional models are far from plateauing, but they still lag behind the efficiency of neural architectures.
- The "Deception" Test: By training on subsets where authors "mimic" the opposite gender (GI A, B subsets), the authors forced the networks to look beyond simple morphological endings. This "adversarial" balancing is what pushes the model to learn deeper stylistic markers.
Figure 1: Illustrates how accuracy scales with training set size for SVM.
Critical Analysis & Conclusion
The study concludes that Model 1 (CNN) is currently the most robust approach for Russian author profiling. The fact that it outperforms LSTMs suggests that for gender identification, local stylistic "shards"—captured by convolutions—are more informative than the long-term temporal dependencies usually captured by LSTMs.
Takeaways
- Traditional ML is the Floor, not the Ceiling: While SVMs are fast, they leave roughly 10-15% of performance on the table compared to Deep Learning.
- Crowdsourcing Works: The success of the
GI cscorpus proves that web-scale data collection via platforms like Amazon Mechanical Turk (or Russian equivalents) is valid for building high-quality forensic NLP tools.
Future Outlook
While the 88% F1-score is impressive, the next frontier is Cross-genre Deception. How well does a model trained on Twitter detect a woman posing as a man in a professional letter? By refining these neural models, the industry moves closer to high-reliability digital forensics and personalized user experience modeling.
