LED: Harmonizing History — Crowdsourcing Musical Reception via Linked Data
Crowdsourcing Linked Data on listening experiences through reuse and enhancement of library data
The article presents the Listening Experience Database (LED), a native Linked Data platform that aggregates over 10,000 subjective accounts of music listening throughout history. It employs a supervised crowdsourcing workflow to reuse and enhance library data (e.g., British Library) and general knowledge bases (DBpedia, MusicBrainz) into a formal RDF graph.
TL;DR
The Listening Experience Database (LED) is a groundbreaking initiative that transforms thousands of subjective historical accounts—found in diaries, letters, and memoirs—into a structured, machine-readable Linked Data graph. By combining supervised crowdsourcing with advanced NLP and semantic mapping, LED allows researchers to query the "who, what, where, and when" of music history with unprecedented precision.
Contextualizing the Listener: The Semantic Gap
In musicology, we often know what was composed (scores) and what was sold (market data), but we rarely have structured data on how individuals actually experienced music. This evidence exists but is trapped in massive, unstructured textual archives. The "semantic gap" here is twofold:
- The Infrastructure Gap: Traditional databases are too rigid for the messy, "fuzzy" nature of historical accounts.
- The Interoperability Gap: Bibliographic data (books) and musical data (compositions) often live in separate silos.
LED bridges these gaps by treating every listening experience as a semantic "Event" that links a Listener, a Piece of Music, a Source Document, and a Time/Place.
Methodology: A Native RDF Approach
Unlike projects that convert a relational database to RDF as a final step, LED is RDF-native. This means the model used for storage is identical to the model used for publication.
The Core Ontological Model
The project reuses established "ontological blocks":
- BIBO & Dublin Core: For representing the complex hierarchy of books, chapters, and specific excerpts.
- Music Ontology: For describing performances, instruments, and genres.
- Event Ontology: Acts as the "glue," connecting the listener to the performance.

Handling Vagueness: The EDTF and NLP NLP Pipeline
One of the most impressive technical feats in LED is its handling of "fuzzy" data. Historical records rarely say "June 6th, 1824 at 14:00." They say "sometime in the 1820s" or "near a village in Italy."
- Temporal: LED implemented a Linked Data version of the Extended Date-Time Format (EDTF), allowing SPARQL queries to handle intervals and approximations.
- Geospatial: Instead of forcing users to pick from a list, LED uses a Supervised Named Entity Extraction pipeline. As a user types a description, a natural language processor identifies potential DBpedia entities, allowing the user to validate the link in real-time.
Critical Results: From Text to Map
The LED dataset now encompasses over 400,000 triples. By interlinking with DBpedia, MusicBrainz, and The British National Bibliography, the project creates a "knowledge web" where a single diary entry in LED can be enriched with metadata about the author's religion from DBpedia and the composer's catalog from MusicBrainz.

The platform's Interactive Geographical Browser demonstrates the power of this approach. It doesn't just show a pin on a map; it leverages mereological relationships (e.g., knowing that "Central Park" is in "Manhattan" which is in the "USA") to allow hierarchical filtering without the crowd ever having to manually enter those parent locations.
Deep Insight & Future Directions
The "magic" of LED lies in its Data Reconciliation workflow. By allowing moderators to merge redundant URIs and map them to authority files like VIAF, the project maintains high data quality while benefiting from the scale of crowdsourcing.
Limitations & Bias
The authors candidly acknowledge a "British bias" in the current dataset (skewed toward 19th-century Britain). This is a common challenge in Digital Humanities, often reflecting the availability of digitized English-language sources rather than the limits of the technology itself.
The Future of Scholarly Crowdsourcing
LED proves that Linked Data is not just a publication format—it is a management paradigm. Moving forward, the project aims to implement web crawlers to automatically detect listening experiences across the web, potentially moving from "supervised entry" to "automated discovery."
Summary
LED is a masterclass in applying Semantic Web technologies to the "messy" reality of the Humanities. It provides a blueprint for how we can turn our shared cultural history into a structured, searchable, and globally connected knowledge base.
