You Can Yak but You Can’t Hide: Breaking Anonymous Social Networks via Virtual Probing
You Can Yak but You Can't Hide: Localizing Anonymous Social Network Users.
This paper presents a novel localization attack on Yik Yak, a popular anonymous social network. By utilizing a "virtual probe" methodology combined with Lasso regression and a centroid heuristic, researchers achieved a Mean Absolute Error (MAE) of only 106 meters and 100% accuracy in identifying the specific college dormitory origin of messages.
TL;DR
Researchers have demonstrated that "anonymity" on location-based social networks like Yik Yak is an illusion. By virtually spoofing GPS locations and using Machine Learning (Lasso Regression), they can localize an anonymous message sender to within 100 meters—enough to pinpoint a specific college dorm room—with 100% accuracy in dormitory classification, all without needing to crack the app's encryption.
The Illusion of Proximity Privacy
Yik Yak became a phenomenon on college campuses by allowing users to post "yaks" visible only to those nearby. Unlike its competitor Whisper, Yik Yak does not show "distance to sender." Most users (and likely the developers) assumed that hiding distance metrics was sufficient to prevent trilateration attacks.
However, the paper identifies a critical Inductive Bias in the system: the app's "user-centric proximity" logic. If moving 50 meters to the north causes a message to disappear from your feed, you have effectively gained a data point about the sender's boundary.
Methodology: The "Remote Stalker" Framework
The genius of this attack lies in its simplicity. Instead of fighting certificate pinning or encrypted traffic (the hard way), the researchers used an "off-the-shelf" automation stack:
- Emulators: Spoofing GPS coordinates to create "virtual probes."
- Sikuli: An automation tool that mimics human "tapping" and "scrolling."
- OCR (FineReader): Converting screenshots of the app into searchable text.
The Layout
Initially, the team used a "honeycomb" layout with 2,880 probes. They discovered that Yik Yak doesn't use a circle to define "nearby"—it uses a square-like shape (vaguely resembling their logo). To optimize, they moved to a Sparse Layout, placing probes in lines extending East, West, North, and South.

Cracking the Code with Machine Learning
The researchers treated localization as a Supervised Learning regression problem.
- Features (): A binary vector indicating whether a specific message was visible at each of the probe locations.
- Labels (): The known (ground truth) longitude and latitude of messages they posted themselves for training.
They utilized Lasso (Least Absolute Shrinkage and Selection Operator) regression. Lasso is particularly effective here because it provides a sparse representation, essentially "selecting" which probes are most representative of the boundary changes.
Experimental Results: Room-Level Precision
The results from the University of California Santa Cruz (UCSC) are startling:
- Accuracy: Mean Absolute Error (MAE) of ~106 meters.
- Contextual Success: They posted yaks from 9 different residential colleges and reached 100% accuracy in identifying which college sent which yak.

The implication is severe: if a professor or a bully can link a message to a specific dorm building, they can cross-reference it with a student directory or class list, reducing the "anonymous" pool of senders from thousands to just one or two individuals.
Critical Insights: Can We Fix This?
The authors argue that typical "obfuscation" (adding noise to coordinates) isn't enough because a dedicated attacker can simply collect more data and "train out" the noise.
The Solution: The authors propose Static Display Regions. Instead of a shifting window that follows the user, the app should use fixed geographic tiles (e.g., "The UCSC Campus" or "The Mission District"). In a static region, every user inside the tile sees the exact same content, making the "feature vector" identical for every probe and rendering the machine-learning attack useless.
Conclusion
This study serves as a masterclass in side-channel attacks. It proves that as long as an application's output (what you see) is dependent on a precise input (where you are), the input can be reverse-engineered. For the "Next-Gen" of anonymous apps, the lesson is clear: true anonymity requires breaking the link between precise geolocation and content delivery.
Limitations
- The attack is computationally expensive (taking hours/days to probe a whole campus with one PC).
- It relies on the OCR being accurate (though modern LLM-based OCR would likely solve the recognition errors mentioned in the paper).
