Boosting Hate Speech Detection: The Power and Pitfalls of Data Augmentation in Portuguese NLP

Publication rights licensed to ACM. ACM acknowledges that this contribution was authored or co-authored by an employee, contractor or affiliate of a national government. As such, the Government retains a nonexclusive, royalty-free right to publish or reproduce this article, or to allow others to do so, for Government purposes only

Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores Data Augmentation (DA) techniques, specifically Random Deletion and machine translation, to enhance Hate Speech Detection in Portuguese social media texts. By applying these methods to LSTM and CNN architectures, the authors achieve improved classification performance on small, scarce datasets.

TL;DR

This research addresses the scarcity of labeled Portuguese hate speech data by testing two Data Augmentation (DA) strategies: Random Deletion and Cross-lingual Translation. While Random Deletion provided a significant performance boost (up to 2.73% accuracy gain) by preventing overfitting, automatic translation from English proved detrimental, highlighting the cultural specificity of offensive language.

Context & Motivation: The Data Desert

Deep learning models like CNNs and LSTMs are "data-hungry." In the realm of hate speech detection, while English datasets are abundant, Portuguese resources are notably small. This leads to a classic bottleneck: overfitting. The authors set out to determine if Computer Vision-style augmentation—artificially expanding the dataset—could be successfully ported to the complex, nuanced world of Natural Language Processing (NLP).

Methodology: Simplicity vs. Complexity

The study utilizes two primary neural architectures:

  1. LSTM: Capturing temporal dependencies in text.
  2. CNN: Utilizing 1D convolutions for local feature extraction (e.g., specific offensive phrases).

For the data, they used a Portuguese dataset of ~5,600 tweets and applied the following DA techniques:

  • Random Deletion (RD): Randomly removing words to create new samples.
  • Cross-lingual Augmentation: Using Google Translate to port English datasets (Waseem & Hovy) into Portuguese.

Model Architectures Figure 1: Comparison of the LSTM (a) and CNN (b) architectures used in the study.

The Core Insight: Why Random Deletion Works

The researchers found that adding "noise" through deletion acts as a powerful regularizer. Interestingly, the biggest gain (2.73%) occurred when they "unfroze" the embedding layer in the CNN. With the augmented data, the model had enough samples to fine-tune word vectors without immediately memorizing the training set.

However, there is a "Goldilocks zone":

  • (Tripling the data): Optimal results.
  • : Performance plummets. Too much deletion destroys the sentence's syntax and meaning, leading the model to learn from "garbage" data.

Results: The Failure of Translation

A critical finding of this paper is the failure of machine-translated data. Despite increasing the volume of training samples, the models performed worse.

Accuracy Comparison Table 1: Performance metrics with and without Data Augmentation (DA).

The authors provide a qualitative analysis of this failure:

  • Linguistic Nuance: Hate speech in English (e.g., "sexist") translates poorly to the informal Portuguese used on Twitter (where "machista" is 25x more common).
  • Cultural Context: Slang and insults are deeply tied to local culture; literal translations often sound unnatural to the classifier, creating a distribution shift that confuses the model.

Critical Analysis & Conclusion

Takeaway

For developers working on low-resource NLP, Random Deletion is a "low-hanging fruit" that offers reliable gains. It allows for more complex models (like CNNs with trainable embeddings) to be used on small datasets without the immediate risk of overfitting.

Limitations

The study focuses on very simple architectures. Modern Transformers (like BERT or RoBERTa) might react differently to Random Deletion. Furthermore, the reliance on automatic translation without manual "localization" remains a significant barrier for cross-lingual transfer learning in toxic speech tasks.

Future Outlook

Future work should explore Back-translation (translating to a second language and back) or Synonym Replacement using Portuguese-specific LLMs, which might preserve the "naturalness" of the text better than simple deletion or naive translation.

Find Similar Papers

Try Our Examples

  • Search for recent studies on Easy Data Augmentation (EDA) specifically applied to low-resource Romance languages for sentiment or toxicity analysis.
  • Which paper first introduced the Easy Data Augmentation (EDA) framework, and how does this study's implementation of Random Deletion differ from the original formulation?
  • Check for research investigating the effectiveness of Back-Translation versus Random Deletion in improving the robustness of social media text classifiers.
Contents
Boosting Hate Speech Detection: The Power and Pitfalls of Data Augmentation in Portuguese NLP
1. TL;DR
2. Context & Motivation: The Data Desert
3. Methodology: Simplicity vs. Complexity
4. The Core Insight: Why Random Deletion Works
5. Results: The Failure of Translation
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations
6.3. Future Outlook