Architecting for the Edge: A Generic Microservice Framework for Environmental Big Data

A Generic Microservice Architecture for Environmental Data Management

2017-01-01
Eric Braun, Thorsten Schlachter, Clemens Düpmeier, Karl-Uwe Stucky, Wolfgang Suess
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a generic microservice architecture designed for managing large-scale, heterogeneous environmental data. By decoupling data types (Master Data, Time Series, Digital Assets) into specialized services and utilizing a polyglot persistence model, the system achieves high horizontal scalability and modularity.

TL;DR

The transition from monolithic Environmental Information Systems (EIS) to scalable cloud-native architectures is no longer optional. This paper presents a generic microservice-based backend that solves the "Big Data" challenge in environmental monitoring. By isolating Master Data, Time Series, and Digital Assets into discrete services and leveraging a polyglot persistence model, the researchers at KIT have created a future-proof blueprint for managing heterogeneous sensor data with high horizontal scalability.

The Monolith Problem: Why Traditional EIS Fail

Historically, environmental data was managed via centralized relational databases (RDBMS). While reliable, this "One Size Fits All" approach faces three critical failures in the era of IoT and Crowdsourcing:

  1. Scaling Bottlenecks: Relational models struggle to scale horizontally when dealing with millions of sensor readings per second.
  2. Rigid Schemas: Environmental data is inherently diverse. Forcing time series, geospatial data, and unstructured PDFs into a single schema leads to fragile codebases.
  3. High Coupling: Changing one part of the data model often requires redeploying the entire monolithic application.

Methodology: The Polyglot & Semantic Solution

The core of the KIT architecture lies in its "Inner Architecture," which separates concerns not just by function, but by data characteristic.

1. The Polyglot Data Model

Instead of a single database, the system uses a Persistence Adapter Layer. For instance:

  • Master Data: Handled via MongoDB (Document-oriented) for flexible JSON schemas.
  • Time Series: Handled via OpenTSDB (Optimized for timestamps).
  • Relationships: Managed by a Link Service backed by Neo4j (Graph database).

2. Architecture Overview

The request flow is managed by an API Gateway that provides a harmonized URL space, masking the complexity of the distributed services from the client.

System Architecture Figure 1: The microservice reference architecture inspired by Gartner, highlighting management and operational capabilities.

3. Semantic Decoupling

A standout feature is the Schema Service. It allows developers to define data structures (like an "Air Measurement Station") in JSON Schema. This means the backend services remain "generic"—they don't need to know what they are storing; they simply validate it against the schema provided at runtime.

Data Service Detail Figure 2: Interaction between Master Data, Time Series, and the Link Service to aggregate complex objects.

Real-World Validation: From Air Quality to Energy

The researchers deployed two distinct prototypes to prove "genericity":

  • Umweltnavigator Bayern: A public-facing map showing air quality. The frontend aggregates data in parallel—fetching the station location from a Geo-service and the pollutant levels from the Time Series service simultaneously.
  • Smart Building Dashboard: A system where dashboard managers can dynamically add widgets. Because of the Schema Service, the dashboard "knows" what data points are available for a new building without any backend code changes.

Prototype Evaluation Figure 3: Graphical user interface of the "Umweltnavigator Bayern" prototype.

Critical Insight: The "Adapter" Advantage

The most valuable takeaway for system architects is the Data Discovery Service and the Adapter pattern. By inserting an abstraction between the Microservice and the Database (e.g., the Time Series Service talks to an OpenTSDB Adapter), the underlying storage technology can be swapped out (e.g., from OpenTSDB to InfluxDB) without altering a single line of business logic. This is the definition of a truly "decoupled" system.

Conclusion and Future Outlook

While the current prototype relies on synchronous REST communication, the authors identify a shift toward asynchronous messaging (Kafka/RabbitMQ) to further increase elasticity. This architecture proves that by embracing microservices and polyglot persistence, we can build environmental systems that are not only scalable but also flexible enough to handle the unpredictable data types of the future IoT landscape.

Key Takeaways for Readers:

  • Decouple by Data Type: Don't just split services by feature; split them by the nature of the data (static vs. streaming).
  • Metadata is King: Use services like a Schema Service to keep your logic generic and your applications "smart."
  • Abstraction Pays Off: Use adapters to prevent database vendor lock-in.

Find Similar Papers

Try Our Examples

  • Search for recent papers that extend microservice-based environmental information systems using Service Mesh technologies like Istio for advanced traffic management.
  • What are the state-of-the-art methods for integrating Semantic Web and Linked Data principles within a containerized microservices environment to improve data discoverability?
  • Investigate how modern Time Series Databases like InfluxDB or TimescaleDB compare to the OpenTSDB adapter approach used in this architectural framework for high-throughput IoT scenarios.
Contents
Architecting for the Edge: A Generic Microservice Framework for Environmental Big Data
1. TL;DR
2. The Monolith Problem: Why Traditional EIS Fail
3. Methodology: The Polyglot & Semantic Solution
3.1. 1. The Polyglot Data Model
3.2. 2. Architecture Overview
3.3. 3. Semantic Decoupling
4. Real-World Validation: From Air Quality to Energy
5. Critical Insight: The "Adapter" Advantage
6. Conclusion and Future Outlook