ATT-BLSTM: Capturing Semantic "Signals" for Deep Gender Identification

Document Model with Attention Bidirectional Recurrent Network for Gender Identification

2019-01-01
Bassem Bsir, Mounir Zrigui
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces an Attention-based Bidirectional Long Short-Term Memory (ATT-BLSTM) network specifically designed for Gender Identification in author profiling. By integrating an attention mechanism with bidirectional recurrent structures, the model effectively identifies key semantic segments within social media texts (Twitter/Facebook) to achieve state-of-the-art accuracy.

TL;DR

Profiling an author’s gender through short, noisy social media posts is a classic NLP challenge. This paper presents a Document Model with Attention Bidirectional Recurrent Networks (ATT-BLSTM). By using a bidirectional LSTM to understand context and an attention layer to highlight key "gender-telling" words, the model achieves a robust 80.23% accuracy on the competitive PAN@CLEF 2018 dataset.

Background: Beyond Hand-Coded Stylometry

Historically, identifying author traits (Gender, Age, Location) required researchers to manually define "stylometric features"—things like the frequency of emojis, punctuation patterns, or specific function words. However, as social media language evolves, these static rules break. Deep Learning offers a way to learn these features automatically, but standard LSTMs suffer from a "dilution" problem: they treat all words with similar importance, often missing the subtle semantic cues that differentiate male and female writing styles.

The Core Insight: Selective Attention

The authors argue that gender is often signaled by specific "key parts" of a sentence rather than the entire sequence.

  • Bidirectionality: By using a BI-LSTM, the model looks at each word in the context of both its preceding and following words, capturing the full "flow" of Twitter-style Arabic.
  • The Attention Filter: Instead of just using the final hidden state of the LSTM, the attention mechanism assigns a weight () to every word. This allows the model to "attend" more to discriminative words while ignoring neutral content.

Methodology: The ATT-BLSTM Architecture

The proposed pipeline consists of four distinct stages:

  1. Embedding Layer: Converts raw Arabic tokens into 300-dimensional dense vectors using Word2Vec (Pre-trained on 4 million Arabic tweets and Wikipedia).
  2. BLSTM Layer: Processes these embeddings in both directions to generate a sequence of hidden states that encode temporal dependencies.
  3. Attention Layer: Computes a score for each hidden state using a trainable attention vector, essentially asking: "How important is this specific word for determining gender?"
  4. Output Layer: A softmax/sigmoid function that classifies the aggregated vector into Male or Female.

Model Architecture Note: The architecture combines the deep sequential learning of LSTMs with the focused weighting of Attention.

Experiments & Real-World Performance

The model was validated on the PAN@CLEF 2018 corpus, which is a gold standard in Author Profiling.

  • Training Setup: 10-fold cross-validation, ADAM optimizer, and Binary Cross Entropy.
  • Dataset Content: Specifically focused on Arabic dialects (Egypt, Gulf, Levantine, Maghrebi), which are notoriously difficult due to linguistic variation.

Results Comparison

Model / AuthorMethodologyAccuracy
Takahashi et al.SVM + TF-IDF81.70% (Baseline)
Bsir & Zrigui (2018)GRU + Stylometric79.00%
This WorkATT-BLSTM80.23%

While traditional SVMs remain highly competitive in author profiling, the ATT-BLSTM shows that neural models can reach comparable performance without the need for the manually-crafted TF-IDF features or complex linguistic dictionaries.

Experiment Results Chart

Critical Analysis & Conclusion

The ATT-BLSTM provides a powerful framework for extracting social identity from text. Its strength lies in its semantic sensitivity—the ability to find the "needle in the haystack" words that indicate gender.

Takeaways for the Industry:

  • Context Matters: For short texts like tweets, bidirectional context is non-negotiable for high performance.
  • Attention is the New Stylometry: Instead of counting commas, we can let attention layers discover the relevant patterns.

Limitations: The model currently focuses primarily on text. As social media becomes more visual, moving toward multimodal profiling (combining images and text) will be the next frontier for reaching the >85% accuracy threshold.

Future Outlook

The authors suggest that this architecture can be easily extended to identify language variety (dialectic detection) and personality traits (the "Big Five"), making it a versatile tool for digital forensics and targeted marketing.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Hierarchical Attention Networks (HAN) specifically for Arabic author profiling tasks beyond gender, such as age or personality traits.
  • Which paper first introduced the specific attention weight calculation formula used in this work (Eq. 7), and how has it been adapted for short-text social media classification?
  • Explore if there are studies applying Transformer-based architectures like BERT or AraBERT to the PAN@CLEF 2018 dataset and how their performance compares to ATT-BLSTM.
Contents
ATT-BLSTM: Capturing Semantic "Signals" for Deep Gender Identification
1. TL;DR
2. Background: Beyond Hand-Coded Stylometry
3. The Core Insight: Selective Attention
4. Methodology: The ATT-BLSTM Architecture
5. Experiments & Real-World Performance
5.1. Results Comparison
6. Critical Analysis & Conclusion
7. Future Outlook