Tour Miner: Transforming Twitter Geodata into Dynamic Travel Insights

A method to discover spots from Twitter for tour miner

2017-11-01
Kazuya Nakahori, Shingo Yamaguchi
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces "Tour Miner," a system designed to automatically discover sightseeing spots by mining geocoded tweets from Twitter. It utilizes a density-based clustering algorithm (DBSCAN) to identify geographical hotspots based on user activities and keyword relevance, transforming raw SNS data into actionable travel insights.

TL;DR

The "Tour Miner" system presents a specialized "Mining Function" that converts the chaos of Twitter's geocoded data into structured sightseeing spots. By applying the DBSCAN clustering algorithm to user timelines filterable by keywords, the system identifies popular locations (spots) that reflect actual human behavior rather than just static map entries.

Background & Motivation: Beyond Static Landmarks

Most existing travel recommendation systems are "dictionary-based"—they use APIs like Google Places to find locations. However, these databases don't tell you how people actually move or how different locations are linked by shared interests.

The authors argue that Social Networking Services (SNS) like Twitter contain a wealth of "living" travel records. The challenge is the "noise": not every tweet with a geocode is a travel spot. The research team’s goal is to build a bridge between raw SNS data and a structured "smelting" process that eventually yields full tour plans.

Methodology: The Two-Step Mining Logic

The authors propose a systematic pipeline to move from a keyword to a geographical cluster.

1. Data Collection (The Crawler)

Using the Twitter API, the system runs a daily script that tracks users posting geocoded tweets in specific regions (e.g., Tokyo). Key to this step is collecting the entire timeline of these users, which provides a sequential history of their movements.

2. Spot Discovery (Spatial Clustering)

This is where the mathematical precision of DBSCAN (Density-Based Spatial Clustering of Applications with Noise) comes in. Unlike K-means, DBSCAN doesn't require you to know the number of clusters in advance.

  • Core Logic: If a certain number of tweets (minPts) are found within a specific radius (ε), that area is marked as a "Spot."
  • Keyword Filtering: The system filters these users by interest keywords (e.g., "shrine") to ensure the discovered spots are relevant to the user's intent.

Overview of our algorithm Fig 1. The two-step process: Scraping user timelines followed by spatial clustering via DBSCAN.

Implementation Architecture

The system is built as a complete web application. The "Collect Tweets" module feeds a database, which the "Mining Module" then processes in real-time based on web-based user queries.

System Configuration Fig 2. Technical architecture of the Tour Miner system, showing the interaction between the Twitter API, the Database, and the Web UI.

Experimental Results: The "Shrine" Test Case

To validate the system, the authors queried "shrine" against a database of roughly 118,000 tweets.

  • Parameters: They set ε = 0.0005 (approx. 50 meters) and minPts = 30.
  • Discovery: The system didn't just find shrines (Meiji, Yasukuni); it also found "associated spots" like Tokyo Sky Tree and Shibuya Station.
  • Insight: This demonstrates that people interested in shrines are also likely to visit major landmarks, providing a more comprehensive "interest-based" map of the city.

Spots discovered by mining Fig 3. Visualization of discovered spots in Tokyo; note how the system maps clusters to real-world landmarks.

Critical Analysis & Conclusion

Takeaway

The Tour Miner approach proves that spatial density of social media activity is a reliable proxy for physical "spots." Its strength lies in its ability to discover spots dynamically based on current trends and specific user interests.

Limitations

A notable limitation mentioned by the authors is the presence of "non-travel" geocoded tweets (e.g., people tweeting from home or work). These act as noise that may create "fake spots."

Future Outlook

The next step for this research is the Smelting Function—taking these discovered spots and the underlying user timelines to reconstruct entire travel routes. If the noise-filtering improves (e.g., using NLP to distinguish between "working at a shrine" and "visiting a shrine"), this could become a powerful tool for automated travel itinerary generation.

Find Similar Papers

Try Our Examples

  • Search for recent studies that use DBSCAN or other density-based clustering algorithms to identify functional urban areas from social media data.
  • What were the original architectural frameworks for the "Tour Miner" system as proposed in the 2016 and 2017 papers by Shingo Yamaguchi et al.?
  • Explore how similar geocoded tweet mining techniques are being applied to real-time event detection or disaster management in smart cities.
Contents
Tour Miner: Transforming Twitter Geodata into Dynamic Travel Insights
1. TL;DR
2. Background & Motivation: Beyond Static Landmarks
3. Methodology: The Two-Step Mining Logic
3.1. 1. Data Collection (The Crawler)
3.2. 2. Spot Discovery (Spatial Clustering)
4. Implementation Architecture
5. Experimental Results: The "Shrine" Test Case
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations
6.3. Future Outlook