Tour Miner: Transforming Twitter Geodata into Dynamic Travel Insights
A method to discover spots from Twitter for tour miner
The paper introduces "Tour Miner," a system designed to automatically discover sightseeing spots by mining geocoded tweets from Twitter. It utilizes a density-based clustering algorithm (DBSCAN) to identify geographical hotspots based on user activities and keyword relevance, transforming raw SNS data into actionable travel insights.
TL;DR
The "Tour Miner" system presents a specialized "Mining Function" that converts the chaos of Twitter's geocoded data into structured sightseeing spots. By applying the DBSCAN clustering algorithm to user timelines filterable by keywords, the system identifies popular locations (spots) that reflect actual human behavior rather than just static map entries.
Background & Motivation: Beyond Static Landmarks
Most existing travel recommendation systems are "dictionary-based"—they use APIs like Google Places to find locations. However, these databases don't tell you how people actually move or how different locations are linked by shared interests.
The authors argue that Social Networking Services (SNS) like Twitter contain a wealth of "living" travel records. The challenge is the "noise": not every tweet with a geocode is a travel spot. The research team’s goal is to build a bridge between raw SNS data and a structured "smelting" process that eventually yields full tour plans.
Methodology: The Two-Step Mining Logic
The authors propose a systematic pipeline to move from a keyword to a geographical cluster.
1. Data Collection (The Crawler)
Using the Twitter API, the system runs a daily script that tracks users posting geocoded tweets in specific regions (e.g., Tokyo). Key to this step is collecting the entire timeline of these users, which provides a sequential history of their movements.
2. Spot Discovery (Spatial Clustering)
This is where the mathematical precision of DBSCAN (Density-Based Spatial Clustering of Applications with Noise) comes in. Unlike K-means, DBSCAN doesn't require you to know the number of clusters in advance.
- Core Logic: If a certain number of tweets (
minPts) are found within a specific radius (ε), that area is marked as a "Spot." - Keyword Filtering: The system filters these users by interest keywords (e.g., "shrine") to ensure the discovered spots are relevant to the user's intent.
Fig 1. The two-step process: Scraping user timelines followed by spatial clustering via DBSCAN.
Implementation Architecture
The system is built as a complete web application. The "Collect Tweets" module feeds a database, which the "Mining Module" then processes in real-time based on web-based user queries.
Fig 2. Technical architecture of the Tour Miner system, showing the interaction between the Twitter API, the Database, and the Web UI.
Experimental Results: The "Shrine" Test Case
To validate the system, the authors queried "shrine" against a database of roughly 118,000 tweets.
- Parameters: They set ε = 0.0005 (approx. 50 meters) and minPts = 30.
- Discovery: The system didn't just find shrines (Meiji, Yasukuni); it also found "associated spots" like Tokyo Sky Tree and Shibuya Station.
- Insight: This demonstrates that people interested in shrines are also likely to visit major landmarks, providing a more comprehensive "interest-based" map of the city.
Fig 3. Visualization of discovered spots in Tokyo; note how the system maps clusters to real-world landmarks.
Critical Analysis & Conclusion
Takeaway
The Tour Miner approach proves that spatial density of social media activity is a reliable proxy for physical "spots." Its strength lies in its ability to discover spots dynamically based on current trends and specific user interests.
Limitations
A notable limitation mentioned by the authors is the presence of "non-travel" geocoded tweets (e.g., people tweeting from home or work). These act as noise that may create "fake spots."
Future Outlook
The next step for this research is the Smelting Function—taking these discovered spots and the underlying user timelines to reconstruct entire travel routes. If the noise-filtering improves (e.g., using NLP to distinguish between "working at a shrine" and "visiting a shrine"), this could become a powerful tool for automated travel itinerary generation.
