BLSTM: Decoding Gender Identity Through Deep Stylometry in Arabic Twitter

Bidirectional LSTM for Author Gender Identification

2018-01-01
Bassem Bsir, Mounir Zrigui
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents a Bidirectional Long Short-Term Memory (BLSTM) approach for author gender identification in Arabic Twitter texts. By combining Word2Vec embeddings with a BLSTM architecture, the authors achieved a 79.23% accuracy, outperforming several deep learning baselines in the PAN@CLEF 2017 competition.

TL;DR

Determining the gender of an anonymous author—known as Author Profiling—is a critical task for marketing and cybersecurity. This paper proposes a Bidirectional Long Short-Term Memory (BLSTM) network coupled with Word2Vec embeddings specifically tuned for the Arabic language. The system achieves an impressive 79.23% accuracy, outperforming previous Recurrent Neural Network (RNN) and Convolutional Neural Network (CNN) architectures on the PAN 2017 dataset.

Background: Why Gender Identification is Hard

Author profiling isn't just about what people say (topic), but how they say it (style). Traditional methods relied on "Bag-of-Words" or Part-of-Speech (POS) n-grams. However, Arabic presents unique challenges:

  • Morphological Richness: A single word can contain a subject, verb, and object.
  • Short Texts: Twitter’s character limit leaves little room for deep stylistic analysis.
  • Non-Linear Style: Stylistic cues are often distributed across a sentence in ways that simple frequency counts (SVMs) or local windows (CNNs) might miss.

Methodology: The Power of Bidirectionality

The core of this work is the Bidirectional LSTM, an architecture designed to "see" the context of a word from both the past (beginning of the tweet) and the future (end of the tweet).

1. Vectorization (Word2Vec)

Before feeding text into the network, words are converted into 300-dimensional vectors. The authors didn't just use a generic model; they trained a skip-gram model on:

  • Arabic Wikipedia: For formal vocabulary.
  • 4 million Tweets: To capture the informal, dialectal nuances of the "Twitter-sphere."

2. The BLSTM Architecture

The model utilizes memory cells with gates (Input, Forget, and Output) to manage information flow. By using a bidirectional approach, the hidden state at any given timestamp contains information from both directions, allowing the model to capture the "global semantic flow" of the tweet.

System Architecture Figure 1: The proposed high-level flow from raw text to gender classification.

Experiments & Performance

The authors evaluated their model against the PAN@CLEF 2017 baseline, specifically focusing on the Arabic subset which includes varieties from the Gulf, Levant, Maghreb, and Egypt.

Key Comparisons:

MethodAccuracy
SVM + tf-idf n-grams (Basile et al.)80.06%
Proposed BLSTM79.23%
CNN (Miura et al.)76.44%
GRU (Kodiyan et al.)71.50%

The results demonstrate that while the SVM with hand-crafted features is still a tough baseline to beat, the BLSTM is significantly more effective than other deep learning variants. It suggests that the sequential, bidirectional nature of LSTMs is better suited for stylometry than the local feature extraction of CNNs.

Accuracy Curve Figure 2: Training and Testing accuracy over 100 epochs, showing stable convergence.

Critical Insight: Sigmoid vs. Others

An interesting finding in the ablation study was the choice of activation units. The authors found that the Sigmoid unit yielded the highest accuracy (78.4% during initial testing) compared to others, which is noteworthy as many modern architectures lean toward ReLU for faster convergence. In specific stylistic classification tasks, the "squashing" 0-1 property of Sigmoid may help in normalizing the varied signals found in social media text.

Conclusion & Future Outlook

The paper confirms that BLSTMs are a powerful tool for Arabic Author Profiling. By capturing the temporal dependencies of writing style, the model bridges the gap between traditional linguistics and modern AI.

Limitations: The model currently focuses only on gender. Future iterations will need to tackle the "joint identification" problem—predicting gender, age, and dialect simultaneously, as these factors are often intertwined in natural language.

Future Work: The logical next step is the integration of Attention Mechanisms to highlight specifically "gendered" keywords and the transition toward Transformer-based pre-trained models (like BERT) to handle the extreme vocabulary sparsity of Arabic dialects.

Find Similar Papers

Try Our Examples

  • Find recent papers on gender identification in Arabic social media that utilize Transformer-based models like AraBERT to compare against LSTM performance.
  • Which study first established the "stylometric" differences between male and female writing styles in forensic linguistics, and how does this paper automate those findings?
  • Investigate if the BLSTM architecture for author profiling has been successfully adapted for dialect identification or detecting social class in multi-dialectal Arabic corpora.
Contents
BLSTM: Decoding Gender Identity Through Deep Stylometry in Arabic Twitter
1. TL;DR
2. Background: Why Gender Identification is Hard
3. Methodology: The Power of Bidirectionality
3.1. 1. Vectorization (Word2Vec)
3.2. 2. The BLSTM Architecture
4. Experiments & Performance
4.1. Key Comparisons:
5. Critical Insight: Sigmoid vs. Others
6. Conclusion & Future Outlook