Tweeting the Future: How Social Media Predicts Your Next Move and Potential Crime

Using Twitter for Next-Place Prediction, with an Application to Crime Prediction

2015-12-01
Mingjun Wang, Matthew S. Gerber
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a text-enriched approach for next-place prediction using Twitter data, specifically integrating TF-IDF textual features into classification and regression models. It achieves a significant breakthrough by correlating these predicted mobility patterns with real-world crime occurrences in Chicago, demonstrating state-of-the-art accuracy in venue-type prediction.

TL;DR

Researchers from the University of Virginia have developed a model that listens to the "implicit hints" in your tweets—like saying "I'm hungry"—to predict where you are going next with over 71% accuracy. More importantly, they've proven that these predicted waves of human movement are statistically linked to where crimes like burglary and theft are likely to strike next.

Background Positioning

While current SOTA models for human mobility focus on "where you've been" (using GPS logs or Foursquare check-ins), this work pivots to "what you're saying." It bridges the gap between Natural Language Processing (NLP) and Spatial-Temporal Criminology, moving from retrospective "hot-spot" mapping to proactive intent-based prediction.

Problem & Motivation: The Silence of the Coordinates

Standard next-place prediction models are "silent"—they treat coordinates as numbers but ignore the human intent. If a user tweets "@joshua: I'm hungry," they aren't at a restaurant yet, but they likely will be soon.

The authors identified two major gaps:

  1. Textual Neglect: Existing models ignore the rich semantic clues in tweets that signal future transitions.
  2. Static Policing: Crime prediction often relies on historical "hot spots," which fail to account for the dynamic flow of potential victims and offenders across the city.

Methodology: The Text-Enriched Engine

The core innovation lies in anchoring ephemeral text to physical venue types (e.g., Food, Nightlife, Residence).

1. Classification & Regression

The authors developed a two-step SVM process:

  • Step 1: Predict if the user will stay or move.
  • Step 2: If moving, use TF-IDF features from the tweet to predict the target Foursquare venue type.

2. The Interaction Effect

Beyond just classifying types, they used Regression to predict the distance to every venue type. They found that an "Interaction Model" (which accounts for the fact that shops and restaurants often cluster together) performed best.

Model Architecture and Process Flow (Note: The paper utilizes a combination of POS tagging via TweetNLP and SVM classification to bridge text with Foursquare venue categories.)

Experiments & Results: Outperforming the Baselines

The results were striking. The Text-Enriched model outperformed standard Markov Chain models by a wide margin.

ModelPrediction Accuracy
Most Frequent Check-in59.0%
Order-2 Markov Model54.7%
Text-Enriched Model (Ours)71.4%

The Crime Correlation

The researchers then mapped Chicago in a 2000-meter grid and found that predicted movements to "Shop & Services" destinations positively correlated with Burglary. Interestingly, "Narcotics" crimes showed a negative correlation with movement toward Universities, likely due to increased security presence in those zones.

Table of Crime Correlations Table showing the significant counts of different crime types used for the correlation analysis.

Critical Analysis & Conclusion

Takeaway

This research proves that "Next-Place Prediction" isn't just for targeted advertising—it's a critical component of "Routine Activity Theory" in the 21st century. By knowing where people intend to go, we can predict the shifting "confluence of victims and offenders."

Limitations

  • Data Sparsity: The model requires users to post at least 20 tweets per month, potentially biasing results toward "power users."
  • Simple NLP: Using TF-IDF is effective but misses the nuanced context that modern Large Language Models (LLMs) could provide.

Future Work

The authors suggest that future systems should integrate user relationship networks (who you talk to) and more complex spatial constraints to turn these correlations into a high-precision, automated crime forecasting system.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Deep Learning or Transformers for next-place prediction using combined textual and trajectory data on Twitter.
  • Which original study established the Routine Activity Theory in criminology, and how have subsequent computational models digitized its three core components (offenders, targets, guardians)?
  • Investigate how social media-based movement prediction has been applied to other urban management tasks such as traffic congestion forecasting or disaster response.
Contents
Tweeting the Future: How Social Media Predicts Your Next Move and Potential Crime
1. TL;DR
2. Background Positioning
3. Problem & Motivation: The Silence of the Coordinates
4. Methodology: The Text-Enriched Engine
4.1. 1. Classification & Regression
4.2. 2. The Interaction Effect
5. Experiments & Results: Outperforming the Baselines
5.1. The Crime Correlation
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations
6.3. Future Work