Analyzing #LasTesis: How Topic Models Decode Global Digital Activism
Analyzing #LasTesis Feminist Movement in Twitter Using Topic Models
This paper presents a computational analysis of the #LasTesis feminist movement on Twitter, comparing Latent Dirichlet Allocation (LDA) and Biterm Topic Model (BTM) for event detection. By processing over 627,000 tweets in multiple languages, the study successfully identifies cross-national socio-political events, highlighting BTM's superiority in handling short-text social media data.
TL;DR
This study investigates the viral #LasTesis feminist movement by analyzing over 620,000 tweets using Latent Dirichlet Allocation (LDA) and the Biterm Topic Model (BTM). The researchers demonstrate that advanced topic modeling can automatically map high-intensity social media activity to specific real-world events, such as protests in Chile and Turkey, while revealing a surprisingly high (92%) positive sentiment amidst political turmoil.
The Challenge: Noise vs. Knowledge in 280 Characters
When a social movement like Las Tesis ("A Rapist in Your Path") goes viral, it generates a massive, chaotic stream of data. For computational sociologists, the challenge is two-fold:
- Context Sparsity: Short tweets don't provide enough word co-occurrence data for traditional models like LDA to work effectively.
- Multilingual Evolution: The movement started in Chile (Spanish) but exploded in Turkey, France, and the UK, requiring models that can capture nuances across languages.
The authors argue that understanding these movements requires moving beyond simple keyword counts to Latent Structure Discovery—finding the "hidden" themes that define the discourse.
Methodology: From Words to Insights
The research team employed a comparative strategy to find the most effective way to categorize the movement's digital footprint.
1. LDA vs. BTM
- LDA (Latent Dirichlet Allocation): Treats each tweet as a mixture of topics. Useful but struggles when tweets are too short to establish clear topic distributions.
- BTM (Biterm Topic Model): The "Short-Text Specialist." Instead of looking at individual tweets, BTM looks at biterms—pairs of words that appear together across the whole collection. This solves the sparsity problem.
2. Time-Series Correlation
By mapping tweet volume against time, the researchers identified that social media isn't just "chatter"; it is a mirror. The peaks in December (shown below) directly correspond to arrests in Istanbul and subsequent protests by female politicians in the Turkish parliament.
Figure 1: Dynamics of the #LasTesis movement showing peaks during real-world interventions.
Key Findings: The Geography of Protest
The results show a clear linguistic shift. Initially, the discourse was dominated by Spanish, centered on Chilean schools ("liceo") and the National Stadium. However, as the movement traveled, the vocabulary evolved.
- Spanish Corpus: Focused on the performance art aspects and specific Chilean locations.
- English/Turkish Corpus: Heavily focused on "protest," "police," and "repression," reflecting the more violent responses in Istanbul.
Figure 2: Word cloud of English tweets highlighting "Turkish," "Police," and "Protest."
Qualitative Strength of BTM
The BTM model provided much "cleaner" clusters. For example, BTM Topic 1 in English clearly grouped words like parliament, protest, Turkish, and ministers, capturing the specific event where Turkish women MPs sang the protest song to defy the government.
Figure 3: Cohesive topics generated by the Biterm Topic Model.
Critical Insight & Conclusion
The study proves that BTM is significantly more robust than LDA for social media analysis due to its ability to aggregate global word co-occurrences. While the sentiment was overwhelmingly positive, the study also highlights the "double-edged sword" of social networks: they allow for rapid global solidarity but also provide a platform where state repression (as seen in the Turkish case) becomes a central theme of the discourse.
Future Outlook: The authors suggest that the next step is integrating Machine Learning classifiers that don't rely on pre-defined dictionaries, allowing for even more nuanced detection of sarcasm or evolving slang within social movements.
