Deep Learning for Social Good: Tackling Gender Violence in Mexican Tweets

Deep Neural Network to Detect Gender Violence on Mexican Tweets

2021-01-01
Grisel Miranda, Roberto Alejo, Carlos Castorena, Eréndira Rendón, Javier Illescas, Vicente García
Summary
Problem
Method
Results
Takeaways
Abstract

This paper develops an enhanced Deep Learning Multilayer Perceptron (DL-MLP) to detect Gender-Based Violence (GBV) in Mexican Spanish tweets. By leveraging CountVectorizer and Random Over Sampling (ROS), the model achieves an Area Under the ROC (AUC) of approximately 90%, significantly outperforming previous baselines in classifying violent vs. non-violent content.

    ## TL;DR
    Researchers have developed a deep learning framework specifically designed to identify Gender-Based Violence (GBV) within the Mexican Twitter landscape. By combining a deep Multilayer Perceptron (MLP) with robust oversampling techniques, the system achieves a strong **90% AUC**, providing a scalable tool for digital safety in regional contexts.

    ## The Motivation: A Pandemic within a Pandemic
    During the COVID-19 lockdowns, Gender-Based Violence didn't just stay behind closed doors; it migrated online. Platforms like Twitter became breeding grounds for harassment. However, detecting this in Mexican Spanish is difficult because:
    *   **Class Imbalance**: Violent tweets are "needles in a haystack" compared to everyday conversation.
    *   **Linguistic Nuance**: Slang, regionalisms, and technical "stop words" in Mexican Spanish require specific preprocessing.

    ## Methodology: Deepening the Architecture
    The core of this research is an evolution of a previous Multilayer Perceptron (MLP) model. While the predecessor used a shallow structure, this work introduces a **7-layer deep architecture** with a trial-and-error optimized neuron count (ranging from 20 to 40 per layer).

    ### Key Technical Pillars:
    1.  **Preprocessing**: A surgical approach to removing hashtags, emojis, and URLs to focus on the semantic weight of the text.
    2.  **Feature Extraction**: Utilizing **CountVectorizer** to transform qualitative tweets into quantitative vectors, resulting in over 27,000 distinct features.
    3.  **Balancing the Scales**: The authors compared **Random Over Sampling (ROS)** and **SMOTE** to ensure the neural network didn't simply ignore the minority "violent" class.

    ![Model Pipeline/Architecture](https://cdn.atominnolab.com/wisdoc/images/20260605-26523122-57df-4829-83c6-8018cf4900fb/page_000_block_000.png)

    ## Experimental Results
    The shift from a shallow to a deep network, combined with better sampling, yielded impressive gains. The **ROS (Random Over Sampling)** method proved superior, maintaining high specificity (detecting non-violent tweets accurately) while dramatically increasing sensitivity to violent ones.

    | Sampling Method | AUC (Current Work) | AUC (Prior Best [14]) | Improvement |
    | :--- | :--- | :--- | :--- |
    | **ROS** | **~89.9%** | 80.8% | **+9.1%** |
    | **SMOTE** | **~88.3%** | 81.1% | **+7.2%** |

    ![Experimental Results Table](https://cdn.atominnolab.com/wisdoc/tables/20260605-26523122-57df-4829-83c6-8018cf4900fb/page_006_block_003.png)

    ## Critical Insights & Future Outlook
    The study reveals a classic trade-off: as the model becomes better at spotting violence (Sensitivity), it occasionally misidentifies non-violent tweets (a slight drop in Specificity). This is a common hurdle in **Class Imbalance** problems.

    **The Road Ahead:**
    While the MLP model is highly efficient, the authors acknowledge that the next frontier involves **Transformers (like BERT)** and **Continuous Learning Systems**. These would allow the model to adapt to the ever-evolving slang and tactics used in online harassment.

    Ultimately, this work serves as a vital benchmark for regional social computing, proving that deep learning is not just for commercial optimization, but a critical tool for human rights protection.

Find Similar Papers

Try Our Examples

  • Search for recent papers using Transformer-based models like BERT or RoBERTa specifically fine-tuned for Mexican Spanish hate speech detection.
  • Which study first introduced the Synthetic Minority Over-sampling Technique (SMOTE), and how have recent variants improved upon it for high-dimensional text data?
  • What are the ethical implications and current SOTA methods for real-time automated moderation of Gender-Based Violence on large-scale social media platforms?
Contents
Deep Learning for Social Good: Tackling Gender Violence in Mexican Tweets
1. TL;DR
2. The Motivation: A Pandemic within a Pandemic
3. Methodology: Deepening the Architecture
3.1. Key Technical Pillars:
4. Experimental Results
5. Critical Insights & Future Outlook