The Force is Automated: Dissecting the 350,000-Account 'Star Wars' Botnet on Twitter

Discovery, Retrieval, and Analysis of the 'Star Wars' Botnet in Twitter

2017-07-31
Juan Echeverría, Shi Zhou
Summary
Problem
Method
Results
Takeaways
Abstract

This paper discovers and analyzes the 'Star Wars' botnet on Twitter, a massive network of over 350,000 accounts that exclusively tweet quotations from Star Wars novels. Using geographic anomalies and textual classification, the authors provide a comprehensive breakdown of the botnet's lifecycle, architecture, and survival strategies.

TL;DR

Researchers at University College London discovered a hidden army of 350,000 Twitter bots that spent years tweeting nothing but random quotes from Star Wars novels. These bots survived for over four years by staying "low-profile"—avoiding spam links, mentions, and high-frequency tweeting—providing a masterclass in how botnets can circumvent modern detection algorithms.

Background: The Ghost in the Machine

In the arms race between Twitter's security teams and botmasters, the general assumption is that bots are loud. We expect them to scream pharmaceutical links, flood political hashtags, or follow celebrities en masse. However, the "Star Wars" botnet proved that silence, or rather, subtle imitation, is a far more effective survival strategy. These accounts were created in mid-2013 and remained largely undetected until this study, primarily because they didn't "look" like bots to a standard algorithm.

The "Unconventional" Discovery

The researchers didn't set out to find a botnet; they found it by looking at a map. When plotting the locations of a 1% random sample of tweets, they noticed something physically impossible.

The Geographic Smoking Gun

While human activity follows population centers, these tweets formed two perfect, low-density "blankets" over North America and Europe. The borders of these blankets were perfectly straight lines parallel to latitude and longitude—a clear signature of a computer script generating random coordinates within a bounding box.

Geographic distribution of bot tweets

Fig 1: The fake geographical distribution shows two distinct rectangles that ignore physical geography (seas, deserts).

Methodology: Identifying the Empire

Once the researchers looked at the text of these "rectangular" tweets, the pattern became obvious: they were snippets of Star Wars novels. Phrases like "Luke’s answer was to put on an extra burst of speed" were being broadcasted by thousands of accounts.

Anatomy of a Star Wars Bot

  • Source: 100% used "Twitter for Windows Phone."
  • Content: Random quotes, sometimes with broken words or hashtags like #teamfollowback.
  • Activity: Extremely low (maximum 11 tweets in a lifetime).
  • Social Graph: No more than 10 followers and 31 friends, mostly following other bots within the same net.

To capture the entire network, the authors used a Naive Bayes Classifier. Because the bots used a very narrow vocabulary (limited to specific novels), the classifier was able to achieve a staggering 99.8% accuracy.

Naïve Bayes Confusion Matrix

Table 1: The high accuracy highlights that while bots mimic humans, they do so within a very rigid, machine-defined "style" that is easily caught once identified.

Why the Hidden Empire Matters

You might ask: If they just tweet old book quotes, why are they dangerous?

  1. Dormant Threats: A botnet of 350,000 accounts is a "sleeping giant." At any moment, the botmaster could pivot from Star Wars quotes to political disinformation or malicious links.
  2. The "Aged" Account Premium: Accounts created in 2013 are highly valuable on the black market. They have "seniority," making them less likely to be flagged by automated spam filters than an account created yesterday.
  3. Experimental Ground Truth: This dataset is a rare "clean" sample. Usually, bot datasets are messy and include many different types of bots. This is a massive, homogeneous group from a single creator.

In-degree and Out-degree distribution

Fig 2: Connectivity analysis shows that these bots exist in a "walled garden," mainly following each other to boost perceived social standing.

Critical Insight: The Failure of Heuristics

The Star Wars botnet demonstrates a critical flaw in current detection: Heuristic avoidance.

  • They don't use URLs (avoids spam filters).
  • They don't mention/reply (avoids interaction filters).
  • They use diverse, human-written text (avoids Markov-chain detection).

The authors argue that the only reason they caught this botnet was a design mistake by the botmaster: the decision to spoof locations in a way that looked suspicious on a map. Had the botmaster simply turned off location services, this army might still be entirely invisible.

Conclusion

The "Star Wars" botnet is a reminder that the most dangerous tools are often the ones trying the hardest not to be noticed. For researchers and platforms, the takeaway is clear: we cannot rely on "common sense" assumptions about what a bot looks like. As bots become more sophisticated, our detection methods must move beyond simple text and profile analysis toward identifying massive, synchronized anomalies in the metadata of the social graph.

Find Similar Papers

Try Our Examples

  • Search for recent papers investigating low-activity or 'sleeper' botnets on social media and their role in long-term influence operations.
  • Which study first identified the use of geographic coordinate spoofing in automated social media accounts, and how has detection of fake GPS data evolved?
  • Analyze research on the commercial market for 'aged' social media accounts and how their perceived 'trustworthiness' is quantified by bot-selling platforms.
Contents
The Force is Automated: Dissecting the 350,000-Account 'Star Wars' Botnet on Twitter
1. TL;DR
2. Background: The Ghost in the Machine
3. The "Unconventional" Discovery
3.1. The Geographic Smoking Gun
4. Methodology: Identifying the Empire
4.1. Anatomy of a Star Wars Bot
5. Why the Hidden Empire Matters
6. Critical Insight: The Failure of Heuristics
7. Conclusion