Deep Learning vs. Traditional ML: Shielding Gender Secrets in Russian Texts

Deep Learning neural nets versus traditional machine learning in gender identification of authors of RusProfiling texts

2018-01-01
Alexander G. Sboev, Ivan Moloshnikov, Dmitry Gudovskikh, Anton Selivanov, Roman B. Rybka, Tatiana Litvinova
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents a comparative study of gender identification in Russian texts using the RusProfiling and RusPersonality corpora, pitting conventional machine learning (SVM, Gradient Boosting) against Deep Learning architectures. The authors introduce "Model 1," a multi-layer Convolutional Neural Network (CNN) that achieves a state-of-the-art F1-score of 88% on the gender profiling task.

TL;DR

In the evolving landscape of stylometry, identifying the gender of an author behind a screen is a classic yet complex task. This paper investigates whether modern Deep Learning can surpass the rigorous baselines set by traditional Machine Learning on Russian corpora. The verdict: Model 1, a deep CNN, dominates with an 88% F1-score, setting a new benchmark for the RusProfiling task and proving that hidden stylistic signatures can be extracted more effectively through deep neural hierarchies than through manual feature engineering.

Problem & Motivation

Russian is a morphologically rich language where gender often manifests clearly through verb endings and adjective suffixes. While this sounds like a "solved" classification problem, the difficulty spikes when:

  • Authors intentionally deceive: Writing as the opposite sex to mask identity.
  • Short, informal texts: Social media (Twitter, FB) lacks the formal structures that traditional SVMs rely on.
  • Domain Adaptation: Models trained on letters often fail when tested on blog posts or reviews.

The authors' intuition was that while traditional ML (SVM/Gradient Boosting) captures "surface" statistics, Deep Learning—specifically CNNs and LSTMs—could capture the rhythm and latent structural patterns of gendered writing that persist even when surface pointers are altered.

Methodology: The Core Architectures

The researchers tested several "data-driven" modeling pipelines. The most significant shift was moving from character n-grams to sophisticated neural topologies.

The CNN Powerhouse (Model 1)

Unlike many NLP tasks that default to Recurrent Neural Networks (RNNs), this paper highlights a deep CNN approach. The architecture utilizes four sequential convolutional layers with ReLU activations followed by Global Max Pooling. This design is specialized for detecting local patterns (like specific phrase turns or punctuation habits) regardless of their position in the text.

The LSTM Contenders (Models 2 & 3)

Stacked Bidirectional LSTMs were deployed to capture long-range dependencies. These models process text both forward and backward, attempting to understand the context surrounding every word or morphological tag.

Model Comparison Logic Note: The study leveraged a diverse set of features, from Word2vec to Morphological vectors (GRM-1/2).

Experiments & SOTA Results

The comparison results were decisive. Throughout various splits of the RusProfiling (Tw, FB, LJ) and RusPersonality corpora, Deep Learning maintained a significant lead.

ModelFeature SetF1-ScoreDelta vs. Baseline
Model 1 (CNN)Seq. feat0.88+38%
Gradient BoostingTF-IDF0.79+29%
SVMChar N-grams0.74+15%

Experimental Results Comparison Table 6: Performance on balanced training sets shows Model 1's clear superiority.

Key Insights from Experiments:

  1. More Data, Better Performance: As seen in the SVM learning curve (Figure 1), traditional models are far from plateauing, but they still lag behind the efficiency of neural architectures.
  2. The "Deception" Test: By training on subsets where authors "mimic" the opposite gender (GI A, B subsets), the authors forced the networks to look beyond simple morphological endings. This "adversarial" balancing is what pushes the model to learn deeper stylistic markers.

SVM Learning Curve Figure 1: Illustrates how accuracy scales with training set size for SVM.

Critical Analysis & Conclusion

The study concludes that Model 1 (CNN) is currently the most robust approach for Russian author profiling. The fact that it outperforms LSTMs suggests that for gender identification, local stylistic "shards"—captured by convolutions—are more informative than the long-term temporal dependencies usually captured by LSTMs.

Takeaways

  • Traditional ML is the Floor, not the Ceiling: While SVMs are fast, they leave roughly 10-15% of performance on the table compared to Deep Learning.
  • Crowdsourcing Works: The success of the GI cs corpus proves that web-scale data collection via platforms like Amazon Mechanical Turk (or Russian equivalents) is valid for building high-quality forensic NLP tools.

Future Outlook

While the 88% F1-score is impressive, the next frontier is Cross-genre Deception. How well does a model trained on Twitter detect a woman posing as a man in a professional letter? By refining these neural models, the industry moves closer to high-reliability digital forensics and personalized user experience modeling.

Find Similar Papers

Try Our Examples

  • Find recent research papers and SOTA methods for gender identification and author profiling in morphologically rich languages like Russian since 2017.
  • Which paper first proposed using Stacked Bidirectional LSTM for authorship attribution, and how does the CNN architecture in this study compare to modern Transformer-based approaches like BERT for Russian stylometry?
  • Explore how the data-driven modeling techniques used in the RusProfiling corpus have been applied to multi-modal author profiling or cross-genre deception detection tasks.
Contents
Deep Learning vs. Traditional ML: Shielding Gender Secrets in Russian Texts
1. TL;DR
2. Problem & Motivation
3. Methodology: The Core Architectures
3.1. The CNN Powerhouse (Model 1)
3.2. The LSTM Contenders (Models 2 & 3)
4. Experiments & SOTA Results
4.1. Key Insights from Experiments:
5. Critical Analysis & Conclusion
5.1. Takeaways
5.2. Future Outlook