Predictive Punctuality: Inferring When Users Visit Locations via Social Network Data

Inferring Time when People Visit a Location Using Social Network Data

2017-12-01
Md. Moniruzzaman, Ken Barker
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a novel approach to infer the check-in time of users in Geo-Social Networks (GSNs) like Foursquare given a known location. By framing time inference as a binary classification task and employing a "Smart Negative Label Generation" (BC-SG) algorithm, the authors achieve a state-of-the-art accuracy of 86%.

TL;DR

While most AI models in Geo-Social Networks (GSNs) try to guess where you are going next, this paper flips the script to predict when you will arrive at a known destination. By engineering a "Smart" negative label generation strategy, the researchers transformed sparse check-in data into a high-accuracy (86%) binary classification task, outperforming traditional random sampling methods by nearly 2x.

Background & Positioning

In the landscape of Geo-Social Networks (Foursquare, Facebook Places, etc.), temporal data is often treated as a secondary feature compared to spatial coordinates. This work positions itself as a specialized temporal inference framework. It moves away from simple average-based arrival time predictions toward a sophisticated machine learning model that understands the "context of absence"—knowing when it is unlikely for a user to visit is just as important as knowing when they might.

The Problem: The Missing "No"

Machine learning models thrive on contrast. However, GSN datasets only record successful "check-ins." We know a user was at a coffee shop at 8:00 AM, but we don't have a record saying they weren't there at 8:00 PM.

Prior works either used simple averages (ignoring the specific day/week) or random sampling. Random sampling often creates "noisy" negative labels—it might accidentally label a likely visit time as a negative one, confusing the model and capping accuracy at around 45%.

Methodology: High-Fidelity Negative Sampling

The core innovation lies in the Smart Negative Label Generation (BC-SG). Instead of choosing random points in time, the algorithm:

  1. Generates candidate negative records by mixing real spatial data with random temporal features.
  2. Uses a Weak Classifier (Naive Bayes) to rank these candidates.
  3. Selects the least probable candidate as the official negative record.

This ensures that the "negative" data points are genuinely distinct from the user's usual behavior, providing a cleaner signal for the final Random Forest classifier.

Overall Distribution of Check-ins Figure: The researchers meticulously analyzed Tier-2 venue types, noting how "Baseball Stadiums" peak in the evening while "Nightclubs" see median activity at 6:00 AM—insights that form the basis of their feature engineering.

Feature Engineering

The model utilizes several "Inductive Biases" inherent to human mobility:

  • disHome: Distance from the user's primary "home" location.
  • Social Score (ss): A normalized metric of how many friends have visited the venue.
  • ctime: Semantic time buckets (Morning, Afternoon, Evening, LateNight).

Experiments & Results

The researchers tested their approach on a massive Foursquare dataset containing over 2.2 million check-ins.

Accuracy Comparison Figure: Comparison of BC-SG (86%) vs. BC-RG (45%) and the N5 Baseline. The smart generation of negative records clearly provides the structural support needed for high-precision inference.

Key Findings:

  • BC-SG Accuracy: 86% — A significant leap for time-series inference.
  • Baseline (N5): Only 36% accuracy when selecting from the next five probable days.
  • Individualized Models: The study found that feature selection varies by user, suggesting that mobility "fingerprints" are highly personal.

Critical Insight & Future Outlook

This paper proves that the "negative label problem" in one-class datasets (like check-ins or purchase histories) can be solved through iterative filtering. The logic is simple: if a weak model is very certain someone won't be there, that's the best data point to teach a strong model the boundaries of behavior.

Limitations: The study requires at least 1000 check-ins per user to achieve these results, which limits its applicability to "power users."

Future Work: The next frontier is "Cold Start" prediction—can we predict the arrival time of a user we've only seen check-in five times? By combining this temporal logic with modern Graph Neural Networks (GNNs), the industry could see even more personalized and non-intrusive recommendation engines.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize synthetic negative sampling or contrastive learning for location-based check-in prediction in Geo-Social Networks.
  • What are the seminal papers regarding Naive Bayes and Random Forest applications in spatio-temporal time-series forecasting, and how has the "Smart Sampling" concept evolved since 2011?
  • Explore research that applies check-in time inference techniques to urban planning, specifically for predicting crowd density and traffic flow in smart cities.
Contents
Predictive Punctuality: Inferring When Users Visit Locations via Social Network Data
1. TL;DR
2. Background & Positioning
3. The Problem: The Missing "No"
4. Methodology: High-Fidelity Negative Sampling
4.1. Feature Engineering
5. Experiments & Results
6. Critical Insight & Future Outlook