GIVA: Unifying Smart City Silos through Ontology-Based Instance Matching

Ontology-based Instance Matching for Geospatial Urban Data Integration

2017-11-07
Vivek R. Shivaprabhu, Booma Sowkarthiga Balasubramani, Isabel F. Cruz
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces GIVA, a geospatial data integration framework that employs ontology-based instance matching to unify disparate urban datasets. Specifically, it demonstrates the integration of "Business Licenses" and "Food Inspections" from the City of Chicago to create a unified view of business entities, supporting predictive analytics and situational awareness.

TL;DR

Researchers at the University of Illinois at Chicago have developed GIVA, a framework that solves a major "Smart City" headache: how to tell if two records in different government databases refer to the same physical business when they don't share a unique ID. By combining spatial ontologies with advanced string-matching algorithms, they unified thousands of Chicago's business records with over 96% accuracy.

The Problem: A Tale of Two (or Ten) Datasets

Imagine you are a city health inspector in Chicago. You have one list for Business Licenses and another for Food Inspections. To predict which restaurant might have a health violation, you need to see both. However, "Seven Eleven" in the license database might be "7-11 Inc." in the inspection log. Without a shared primary key (like a Social Security number for businesses), these records remain "information silos."

The challenges are fourfold:

  1. Data Heterogeneity: Different formats (CSV, OWL, XML) and naming conventions.
  2. Representational Variations: Typos, abbreviations, and different address formats for the same building.
  3. One-to-Many Relationships: One business entity might have dozens of annual licenses or inspection reports.
  4. Spatial Ambiguity: Co-located businesses (multiple shops in one mall) make location-only matching impossible.

Methodology: The GIVA Framework

GIVA (Geospatial data Integration, Visualization, and Analytics) addresses these challenges through a modular, four-phase pipeline.

1. The Strategy: Ontologies and Blocking

To prevent the system from trying to match every single business in Chicago against every other business (which would be computationally impossible), GIVA uses Blocking. It employs a Spatial Ontology (Country -> State -> City -> Zip) to ensure it only compares businesses within the same geographic area. It also uses a Business Ontology to avoid comparing a "Daycare Center" with a "Liquor Store."

2. The Core: Instance Matching Logic

Rather than relying on a single algorithm, GIVA generates a similarity vector. It calculates several scores for Business Names and Addresses, including:

  • ISub Similarity: A metric specifically designed for ontology matching that is more robust than standard Levenshtein distance.
  • Weighted Jaccard: Measures the overlap of word components.
  • String Cleaning: Removing special characters and normalizing cases.

GIVA System Architecture

The system uses a high threshold (90% for names, 95% for addresses) to prioritize Precision—ensuring that when the system says two records are a match, they almost certainly are.

Experimental Results: Cleaning Up Chicago's Data

The team tested GIVA on the Chicago ZIP code 60606. In this single area:

  • Internal De-duplication: 16,853 business licenses were collapsed into 3,522 unique business entities.
  • Cross-Dataset Matching: 2,179 food inspections were processed, and GIVA successfully matched 96.7% of them to the correct business license record.

Sample Match Results Table: Examples of matches and blocked pairs based on similarity scores.

Why It Matters: From Data to Action

The output of GIVA is a Unified Business Model served via a RESTful API. This doesn't just clean up spreadsheets; it changes how the city operates.

By integrating this into platforms like OpenGrid, city administrators can see a single "pin" on a map that contains the entire history of a location—licenses, violations, and even Yelp reviews. This visual and analytical unity allows for Predictive Analytics: inspectors can now identify patterns of violations across the city and intervene before a public health crisis occurs.

OpenGrid Visualization OpenGrid UI displaying unified data from multiple sources for a single restaurant.

Conclusion and Future Outlook

GIVA proves that semantic web technologies are not just theoretical; they are practical tools for urban management. The authors plan to expand the system to include Google Places and Yelp data, which will bring the "Voice of the Citizen" into official government workflows. While external validation of third-party data remains a challenge (due to the lack of official keys), GIVA’s modular architecture provides a scalable path forward for the next generation of Smart Cities.

Find Similar Papers

Try Our Examples

  • Search for recent papers using AgreementMakerLight (AML) or similar ontology matching tools for real-time geospatial data integration in smart cities.
  • Which original research proposed the ISub similarity metric, and how has it been evolved for handling noisy urban data in more recent Entity Resolution studies?
  • Examine how the GIVA framework's instance matching methodology has been extended to incorporate unstructured social media data like Yelp reviews for predictive urban analytics.
Contents
GIVA: Unifying Smart City Silos through Ontology-Based Instance Matching
1. TL;DR
2. The Problem: A Tale of Two (or Ten) Datasets
3. Methodology: The GIVA Framework
3.1. 1. The Strategy: Ontologies and Blocking
3.2. 2. The Core: Instance Matching Logic
4. Experimental Results: Cleaning Up Chicago's Data
5. Why It Matters: From Data to Action
6. Conclusion and Future Outlook