Analyzing #LasTesis: How Topic Models Decode Global Digital Activism

Analyzing #LasTesis Feminist Movement in Twitter Using Topic Models

2020-01-01
Sebastián Rodríguez, Héctor Allende-Cid, Cristian González, Rodrigo Alfaro, Claudio Elortegui, Wenceslao Palma, Pedro Santander
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a computational analysis of the #LasTesis feminist movement on Twitter, comparing Latent Dirichlet Allocation (LDA) and Biterm Topic Model (BTM) for event detection. By processing over 627,000 tweets in multiple languages, the study successfully identifies cross-national socio-political events, highlighting BTM's superiority in handling short-text social media data.

TL;DR

This study investigates the viral #LasTesis feminist movement by analyzing over 620,000 tweets using Latent Dirichlet Allocation (LDA) and the Biterm Topic Model (BTM). The researchers demonstrate that advanced topic modeling can automatically map high-intensity social media activity to specific real-world events, such as protests in Chile and Turkey, while revealing a surprisingly high (92%) positive sentiment amidst political turmoil.

The Challenge: Noise vs. Knowledge in 280 Characters

When a social movement like Las Tesis ("A Rapist in Your Path") goes viral, it generates a massive, chaotic stream of data. For computational sociologists, the challenge is two-fold:

  1. Context Sparsity: Short tweets don't provide enough word co-occurrence data for traditional models like LDA to work effectively.
  2. Multilingual Evolution: The movement started in Chile (Spanish) but exploded in Turkey, France, and the UK, requiring models that can capture nuances across languages.

The authors argue that understanding these movements requires moving beyond simple keyword counts to Latent Structure Discovery—finding the "hidden" themes that define the discourse.

Methodology: From Words to Insights

The research team employed a comparative strategy to find the most effective way to categorize the movement's digital footprint.

1. LDA vs. BTM

  • LDA (Latent Dirichlet Allocation): Treats each tweet as a mixture of topics. Useful but struggles when tweets are too short to establish clear topic distributions.
  • BTM (Biterm Topic Model): The "Short-Text Specialist." Instead of looking at individual tweets, BTM looks at biterms—pairs of words that appear together across the whole collection. This solves the sparsity problem.

2. Time-Series Correlation

By mapping tweet volume against time, the researchers identified that social media isn't just "chatter"; it is a mirror. The peaks in December (shown below) directly correspond to arrests in Istanbul and subsequent protests by female politicians in the Turkish parliament.

Time series of original messages and retweets Figure 1: Dynamics of the #LasTesis movement showing peaks during real-world interventions.

Key Findings: The Geography of Protest

The results show a clear linguistic shift. Initially, the discourse was dominated by Spanish, centered on Chilean schools ("liceo") and the National Stadium. However, as the movement traveled, the vocabulary evolved.

  • Spanish Corpus: Focused on the performance art aspects and specific Chilean locations.
  • English/Turkish Corpus: Heavily focused on "protest," "police," and "repression," reflecting the more violent responses in Istanbul.

Word Cloud Analysis Figure 2: Word cloud of English tweets highlighting "Turkish," "Police," and "Protest."

Qualitative Strength of BTM

The BTM model provided much "cleaner" clusters. For example, BTM Topic 1 in English clearly grouped words like parliament, protest, Turkish, and ministers, capturing the specific event where Turkish women MPs sang the protest song to defy the government.

BTM Topic Results Figure 3: Cohesive topics generated by the Biterm Topic Model.

Critical Insight & Conclusion

The study proves that BTM is significantly more robust than LDA for social media analysis due to its ability to aggregate global word co-occurrences. While the sentiment was overwhelmingly positive, the study also highlights the "double-edged sword" of social networks: they allow for rapid global solidarity but also provide a platform where state repression (as seen in the Turkish case) becomes a central theme of the discourse.

Future Outlook: The authors suggest that the next step is integrating Machine Learning classifiers that don't rely on pre-defined dictionaries, allowing for even more nuanced detection of sarcasm or evolving slang within social movements.

Find Similar Papers

Try Our Examples

  • Search for recent studies comparing Biterm Topic Model (BTM) and BERT-based dynamic topic modeling for social media event detection.
  • Which paper originally proposed the Biterm Topic Model, and how does its objective function differ from the standard Latent Dirichlet Allocation?
  • Have there been applications of BTM or multi-lingual topic models in analyzing the "Me Too" movement or other large-scale human rights protests on X (formerly Twitter)?
Contents
Analyzing #LasTesis: How Topic Models Decode Global Digital Activism
1. TL;DR
2. The Challenge: Noise vs. Knowledge in 280 Characters
3. Methodology: From Words to Insights
3.1. 1. LDA vs. BTM
3.2. 2. Time-Series Correlation
4. Key Findings: The Geography of Protest
4.1. Qualitative Strength of BTM
5. Critical Insight & Conclusion