ADMIRE: Bridging the Gap in Environmental Data Integration through Distributed Mining
Application of ADMIRE Data Mining and Integration Technologies in Environmental Scenarios
This paper introduces the Data Integration Engine for Environmental Data (DIEED), a core component of the EU FP7 project ADMIRE. It utilizes the Data Mining and Integration Language (DMIL) and OGSA-DAI framework to unify distributed, heterogeneous environmental data sources for predictive tasks like flood forecasting and rainfall movement.
TL;DR
The ADMIRE project shifts the paradigm of environmental forecasting from resource-heavy physical modeling to agile, data-driven mining. By introducing the Data Integration Engine for Environmental Data (DIEED), researchers can now orchestrate complex workflows across distributed, heterogeneous data sources—such as GRIB files and SQL databases—using a unified language (DMIL) to predict floods and moisture movement with higher resilience.
1. The Challenge: Data Silos vs. Natural Disasters
Environmental phenomena, particularly floods, do not respect organizational boundaries. However, the data required to predict them accurately—rainfall, soil saturation, reservoir levels—are often "jailed" within different institutions (e.g., SHMI and SWE in Slovakia).
Prior work relied heavily on Physical Modeling (e.g., FESWMS), which presents two major flaws:
- Data Rigidity: These models usually require a complete set of input variables to run; if one sensor fails, the model crashes.
- Interoperability Hell: Combining gridded binary (GRIB) meteorological output with relational database hydrological data requires bespoke, non-reusable scripts.
2. Methodology: From Physical Laws to Data Insights
The core philosophy of the ADMIRE engine is to treat data integration as a workflow of reusable components rather than a static ETL (Extract, Transform, Load) process.
The APE Architecture
The authors introduce Application Processing Elements (APEs). Think of these as modular "integration blocks" that can be executed at the source of the data to minimize heavy data transfers.
- Logic Extraction: Using DMIL (Data Mining and Integration Language), the system provides an implementation-independent way to describe how data should be cleaned and merged.
- Decentralized Execution: By leveraging OGSA-DAI, the engine submits sub-workflows directly to remote servers, transforming raw, heterogeneous formats into a common format (lists of tuples) before they ever hit the central integration node.
Figure 1: The APE workflow for the ORAVA scenario, showing the orchestration of remote data retrieval and local transformation.
3. Case Study: The ORAVA Scenario
The effectiveness of the engine was tested on the ORAVA reservoir scenario in Slovakia. The goal: predict water height and temperature propagation downstream.
The Heterogeneity Puzzle:
- Predictors: Rainfall/Temperature (GRIB files), Reservoir Discharge (Relational DB).
- Targets: Hydrological station measurements (Relational DB).
The system successfully synchronized these disparate streams into a unified temporal matrix (as seen in the table below), allowing data mining algorithms to define correlations that physical models might overlook.
Table 1: Integrated data view combining disparate sources into a single mining-ready temporal sequence.
4. Why This Matters: Graceful Degradation
Perhaps the most significant insight from this work is the concept of graceful degradation. In traditional environmental simulation, missing data is a fatal error. In the ADMIRE data-mining approach, the system can still produce a "best-guess" forecast based on historical statistical patterns even when real-time sensor data is partially unavailable.
5. Summary and Future Look
The ADMIRE project demonstrates that the future of environmental science lies in the "Data-Information-Knowledge" pipeline. By abstracting away the complexity of distributed data access through DIEED and DMIL, experts can focus on the mining rather than the plumbing.
Limitations: While the prototype is robust within the Slovakian pilot areas, scaling this to a pan-European level would require broader adoption of the DMIL standard and more intensive computational resources for the complex APE workflows.
Future Work: The transition to the O3 (Ozone) scenario suggests that this framework is versatile enough to move beyond water management into air quality and broader climate change monitoring.
