City-Stories: Resurrecting Urban History through AI and Collective Intelligence
City-Stories: Combining Entity Linking, Multimedia Retrieval, and Crowdsourcing to Make Historical Data Accessible
The paper introduces City-Stories, an integrated system for managing and exploring digitized historical image collections. It leverages a hybrid approach combining Entity Linking (via WikiData), the vitrivr multimedia retrieval engine, and Gamified Crowdsourcing to generate metadata for historical archives that are often unlabeled or poorly documented.
TL;DR
City-Stories is a hybrid platform designed to make "silent" historical archives searchable. It bridges the gap between raw digitized images and structured knowledge by combining WikiData Entity Linking, advanced Multimedia Retrieval (vitrivr), and Crowdsourcing games. It allows users to search history via sketches, locations, or time, transforming passive archives into interactive, community-driven "stories."
The Problem: The "Metadata Void" in Digital Heritage
Digital preservation is more than just scanning old photos. Many historical collections are essentially "dead data" because they lack descriptive metadata. Without categories, tags, or geocoordinates:
- Rare Entities (obscure buildings or long-gone local figures) go unrecognized by standard AI.
- Spatio-temporal localization is nearly impossible for machine learning models when the visual landmarks have changed significantly over decades.
- Heterogeneity across different museum databases prevents researchers from finding related items across collections.
Methodology: A Triple-Engine Architecture
The authors propose a system that operates in both offline (automated) and online (human-in-the-loop) phases.
1. Semantic Data Expansion
Instead of relying solely on exact matches, City-Stories uses Word Embeddings to infer relationships between entities. If an entity is too rare to exist in WikiData, the system localizes the closest embedding in vector space to provide context.
2. Multi-modal Retrieval (vitrivr)
The system integrates the vitrivr engine, enabling sophisticated query modes:
- Query-by-Sketch (QbS): You don't need words; just draw the outline of a cathedral or a bridge.
- Query-by-Location/Time (QbL/QbT): Filters results based on the "where" and "when."

3. Crowdsourcing & Gamification
Recognizing that local citizens are often the best "historical databases," the system includes four specific crowdsourcing tasks:
- Location-Finder & Year-Finder: Pinpointing the origin of an image.
- Annotation-Competition: A competitive game for tagging images to increase engagement.
- Validator: Humans verify the AI's suggested tags, ensuring data quality.
Experiments & Core Components
The system utilizes a modular backend consisting of MongoDB and PostgreSQL for metadata, and CottontailDB (an open-source database optimized for multimedia retrieval) for vector features.
Figure: The data expansion and retrieval pipeline connecting the knowledge base to the end-user interface.
By fusing these sources, the "DigitalValais" and "Médiathèque" datasets become cross-searchable. The integration of RESTful APIs based on the OpenAPI standard ensures that this architecture can be extended to other cultural heritage institutions.
Critical Insight & Conclusion
Takeaway
The genius of City-Stories lies in its Hybrid Intelligence. It acknowledges that while AI can handle the "heavy lifting" of vectorizing millions of images, it lacks the nuanced local knowledge that a resident of a city might have. By turning metadata generation into a gamified social activity, the authors have created a self-sustaining ecosystem for cultural preservation.
Limitations & Future Work
The paper relies heavily on manual participation for the "wisdom of the crowd" to actually manifest. If the user base is small, the metadata growth might stall. Future research should look into Active Learning, where the system identifies which specific images most need human intervention to maximize the efficiency of the crowdsourcing module.
Keywords: Multimedia Retrieval, Entity Linking, Semantic Data, Crowdsourcing, Digital Humanities.
