EventMine: Bridging the Gap Between Unstructured News and Geographic Event Visualization

High‐level event identification in social media

2018-05-18
Zolzaya Dashdorj, Erdenebaatar Altangerel
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces EventMine, an automated system for identifying and geolocating high-level social events from unstructured web sources and social media. It combines Latent Dirichlet Allocation (LDA) for topic modeling with Google Custom Search API for geo-referencing, achieving a baseline identification accuracy of 56%.

TL;DR

EventMine is a specialized pipeline designed to transform chaotic social media feeds and news articles into a structured, map-based event explorer. By integrating Latent Dirichlet Allocation (LDA) for thematic analysis and external Geometric APIs for location tagging, it provides a framework for understanding social dynamics in regions where structured event metadata is severely lacking.

Background & Motivation

In the digital age, we are overwhelmed by information, yet "knowing everything" is harder than ever. While we have plenty of news articles, they lack the spatial and thematic structure needed for advanced recommendation systems. Most prior work assumes the existence of GPS metadata (geo-tagging), but the vast majority of web content is just "free text."

The authors identify a critical gap: How do we take a plain-text news report about a football match or a protest and automatically determine where it is and what it's about? This is particularly difficult for languages like Mongolian, where vocabulary differences and accusative naming conventions create high linguistic noise.

Methodology: The EventMine Architecture

The EventMine system operates through a clean, four-stage pipeline designed to move from raw data to visual insight.

1. Data Harvesting and Parsing

The system uses BeautifulSoup to crawl web sources, extracting a specific event schema: 〈place, title, category, content, date, photo, topic, url〉.

2. The Geolocation Challenge

Since most articles don't include latitude/longitude, the system extracts "place names" from the text and queries the Google Custom Search API. This component is the system's "eye"—mapping textual entities to physical coordinates.

3. Topic Discovery with FastLDA

To categorize events without manual labeling, the authors use Fast LDA. However, raw text is often too noisy for LDA. To solve this, they use the ToPMine approach to extract "representative phrases" first, ensuring the topic model clusters events based on meaningful concepts rather than stop-words.

The architecture of event recognition model, EventMine

Experimental Insights

The researchers tested EventMine on three distinct data sources: Ikon (General news), Ticket (Structured event site), and Facebook.

  • Structured Success: On the "Ticket" source, where data is already somewhat clean, the system reached 85.7% accuracy.
  • Unstructured Struggle: On general news (Ikon), accuracy dropped to 27.4%, highlighting how "noisy" professional news can be when it lacks clear location mentions.
  • The "Sweet Spot" for Topics: By measuring Perplexity (a metric of how well the model predicts the data), the authors found that 7 distinct topics provided the most coherent categorization for their dataset.

The perplexity over train and test set

Critical Analysis & Future Directions

The core achievement of this paper is the creation of a functional prototype that handles the end-to-end "Text-to-Map" pipeline.

The Bottleneck: The heavy reliance on the Google Custom Search API is a double-edged sword. While powerful, it struggles with local Mongolian naming variations. The authors acknowledge that a "richer gazetteer" (a geographic dictionary) specifically for the local region would drastically improve accuracy.

Future Outlook: The next step for this technology is Personalization. By knowing where an event is and what topic it covers (via LDA), the system can begin to "recommend" events to users based on their historical proximity and interests. As NLP models move toward Large Language Models (LLMs), the "Event Recognizer" component could see massive upgrades in its ability to handle linguistic ambiguity.

Takeaway

EventMine demonstrates that even with limited resources and unstructured data, a combination of classical topic modeling and modern mapping APIs can provide a high-level "pulse" of social activities within a city or country.

The application interface of EventMine

Find Similar Papers

Try Our Examples

  • Search for recent papers focusing on event extraction and geolocation in low-resource languages like Mongolian or similar morphologically rich languages.
  • What are the state-of-the-art alternatives to LDA for short-text topic modeling in social media event detection pipelines?
  • Find studies that integrate external gazetteers with Deep Learning models to solve place name disambiguation in unstructured news articles.
Contents
EventMine: Bridging the Gap Between Unstructured News and Geographic Event Visualization
1. TL;DR
2. Background & Motivation
3. Methodology: The EventMine Architecture
3.1. 1. Data Harvesting and Parsing
3.2. 2. The Geolocation Challenge
3.3. 3. Topic Discovery with FastLDA
4. Experimental Insights
5. Critical Analysis & Future Directions
6. Takeaway