Predictive Punctuality: Inferring When Users Visit Locations via Social Network Data
Inferring Time when People Visit a Location Using Social Network Data
This paper introduces a novel approach to infer the check-in time of users in Geo-Social Networks (GSNs) like Foursquare given a known location. By framing time inference as a binary classification task and employing a "Smart Negative Label Generation" (BC-SG) algorithm, the authors achieve a state-of-the-art accuracy of 86%.
TL;DR
While most AI models in Geo-Social Networks (GSNs) try to guess where you are going next, this paper flips the script to predict when you will arrive at a known destination. By engineering a "Smart" negative label generation strategy, the researchers transformed sparse check-in data into a high-accuracy (86%) binary classification task, outperforming traditional random sampling methods by nearly 2x.
Background & Positioning
In the landscape of Geo-Social Networks (Foursquare, Facebook Places, etc.), temporal data is often treated as a secondary feature compared to spatial coordinates. This work positions itself as a specialized temporal inference framework. It moves away from simple average-based arrival time predictions toward a sophisticated machine learning model that understands the "context of absence"—knowing when it is unlikely for a user to visit is just as important as knowing when they might.
The Problem: The Missing "No"
Machine learning models thrive on contrast. However, GSN datasets only record successful "check-ins." We know a user was at a coffee shop at 8:00 AM, but we don't have a record saying they weren't there at 8:00 PM.
Prior works either used simple averages (ignoring the specific day/week) or random sampling. Random sampling often creates "noisy" negative labels—it might accidentally label a likely visit time as a negative one, confusing the model and capping accuracy at around 45%.
Methodology: High-Fidelity Negative Sampling
The core innovation lies in the Smart Negative Label Generation (BC-SG). Instead of choosing random points in time, the algorithm:
- Generates candidate negative records by mixing real spatial data with random temporal features.
- Uses a Weak Classifier (Naive Bayes) to rank these candidates.
- Selects the least probable candidate as the official negative record.
This ensures that the "negative" data points are genuinely distinct from the user's usual behavior, providing a cleaner signal for the final Random Forest classifier.
Figure: The researchers meticulously analyzed Tier-2 venue types, noting how "Baseball Stadiums" peak in the evening while "Nightclubs" see median activity at 6:00 AM—insights that form the basis of their feature engineering.
Feature Engineering
The model utilizes several "Inductive Biases" inherent to human mobility:
- disHome: Distance from the user's primary "home" location.
- Social Score (ss): A normalized metric of how many friends have visited the venue.
- ctime: Semantic time buckets (Morning, Afternoon, Evening, LateNight).
Experiments & Results
The researchers tested their approach on a massive Foursquare dataset containing over 2.2 million check-ins.
Figure: Comparison of BC-SG (86%) vs. BC-RG (45%) and the N5 Baseline. The smart generation of negative records clearly provides the structural support needed for high-precision inference.
Key Findings:
- BC-SG Accuracy: 86% — A significant leap for time-series inference.
- Baseline (N5): Only 36% accuracy when selecting from the next five probable days.
- Individualized Models: The study found that feature selection varies by user, suggesting that mobility "fingerprints" are highly personal.
Critical Insight & Future Outlook
This paper proves that the "negative label problem" in one-class datasets (like check-ins or purchase histories) can be solved through iterative filtering. The logic is simple: if a weak model is very certain someone won't be there, that's the best data point to teach a strong model the boundaries of behavior.
Limitations: The study requires at least 1000 check-ins per user to achieve these results, which limits its applicability to "power users."
Future Work: The next frontier is "Cold Start" prediction—can we predict the arrival time of a user we've only seen check-in five times? By combining this temporal logic with modern Graph Neural Networks (GNNs), the industry could see even more personalized and non-intrusive recommendation engines.
