Bots and Cyborgs: Unmasking Digital Astroturfing in the 2017 Chilean Election
Detection of Bots and Cyborgs in Twitter: A Study on the Chilean Presidential Election in 2017
This study investigates the presence of automated accounts during the 2017 Chilean Presidential Election using a hybrid manual and machine learning approach. By extracting 61 heterogeneous features from tweets and user metadata, the researchers developed classification models (SVM, Random Forest) to distinguish between humans, bots, and "cyborgs" (human-assisted bots), specifically examining how these entities influenced political discourse.
TL;DR
Social media is the new battleground for political influence, but not all soldiers are human. This study analyzes the 2017 Chilean Presidential Election, revealing how "cyborgs"—human-assisted automated accounts—were used to inflate candidate popularity. While machine learning models (SVM, Random Forest) showed promise in controlled training (83% accuracy), their real-world performance dropped to near 56%, highlighting the sophistication of modern political bots.
The "Cyborg" Menace: Beyond Simple Automation
The research identifies a critical evolution in social media manipulation: the Cyborg. Unlike traditional bots that are 100% autonomous, cyborgs are accounts controlled by humans using automation tools like TweetDeck. These accounts can coordinate massive "retweet storms" in milliseconds, granting a candidate false credibility while evading simple bot-detection filters that look for purely "robotic" timing.
The authors' intuition was triggered by a simple observation: during the September 2017 debates, certain candidates saw massive spikes in mentions (over 4,000 in an hour) that consisted almost entirely of retweets with zero original content.
Methodology: The 6-Dimensional Feature Space
To catch these digital ghosts, the researchers extracted 61 unique features categorized into six domains:
- User-based: Profile metadata (followers, account age).
- Friends: Patterns in mentions and retweeting behavior.
- Network: The structure of hashtag and retweet interactions.
- Temporal: The distribution of time intervals between tweets.
- Content/Language: POS-tagging (verbs/adjectives) and entropy of the text.
- Sentiment: The emotional intent (positive/negative) of the messages.
Figure 1: Comparison of candidate mentions during the debate shows abnormal spikes suggesting coordinated automation.
Why the Models Stumbled: The Gap Between Theory and Reality
The researchers tested four major algorithms: Random Forest, AdaBoost, Decision Trees, and Support Vector Machines (SVM).
The Performance Paradox
In the training phase (using a mix of Botometer-labeled and manually tagged data), the results were stellar:
- Accuracy: ~0.83 across all models.
- Precision/Recall: Highly balanced.
However, when applied to the actual Chilean testing set, the performance plummeted:
- SVM Accuracy: 0.56
- Random Forest Accuracy: 0.55
Table 1: The testing results showcase a significant "reality gap" in bot detection.
The Forensic Evidence
The data revealed that candidate Marco EnrÃquez-Ominami had a retweet-to-original-tweet ratio of roughly 6:1, significantly higher than his peers. Furthermore, a high percentage of his support came via TweetDeck, a platform that facilitates multi-account synchronization.
Critical Insight & Conclusion
The study proves that political actors are no longer using "dumb" bots. Instead, they utilize Hybrid/Cyborg systems that carry the "Inductive Bias" of human behavior but the "Scale" of automation.
The failure of standard ML models to maintain high accuracy during testing suggests that:
- Context Matters: Bot behaviors in a Chilean political context differ from the general "Global" bot datasets.
- Feature Evolution: Static features (like followers/friends) are being "gamed" by bot creators. Future research must look deeper into Graph-based representations to map the hidden connections between these accounts.
Takeaway: As we approach more global elections, the ability to distinguish between a "volunteer" and a "coordinated cyborg" remains one of the hardest—and most vital—challenges in computational social science.
