CrODA-gator: Democratizing Crowdsourced Data through Open Access PaaS
CrODA-gator: An Open Access CrowdSourcing Platform as a Service
This paper introduces CrODA-gator, an Open Access Crowdsourcing Platform as a Service (PaaS) designed for scalable sensor data collection and GIS mapping. It bridges the gap in existing platforms by providing public open access to aggregated datasets via well-documented APIs and bulk data import features.
TL;DR
CrODA-gator is a cloud-based Platform as a Service (PaaS) that enables public access to aggregated crowdsourced sensor data. By introducing efficient mapping algorithms (RIE) and robust outlier filtering (RTDP), it provides developers with a scalable infrastructure to collect and utilize real-time data from heterogeneous mobile devices.
Background & Motivation
With over two billion smartphone users worldwide, mobile devices have become opportunistic data generators equipped with over 20 sensors each. Despite this potential, existing platforms like Device Analyzer or OpenStreetMap often suffer from silos—either they are closed-access, restricted to specific domains (navigation only), or lack robust tools for bulk data management.
The authors identify three primary challenges:
- Data Management: Handling massive volumes of heterogeneous data without distortion.
- Integrity: Filtering out malicious or accidental outliers without relying on costly intrinsic metrics.
- Accessibility: Providing a seamless interface for third-party developers to utilize these datasets.
Methodology: The CrODA-gator Architecture
CrODA-gator is built on a modular MVC (Model-View-Controller) architecture deployed on the Microsoft Azure Cloud. It consists of a dedicated Android application for sensing, a cloud-based data store, and a suite of Open APIs.
1. Recursive Indexing with Elimination (RIE)
To visualize data on a map, the platform must aggregate thousands of points into meaningful grids. The RIE algorithm optimizes this by projecting geographic coordinates to screen grids and, crucially, removing data points from the processing set once they are assigned to a grid. This prevents redundant checks and significantly boosts performance as the viewport scales.

2. Relative Threshold Divergence Purge (RTDP)
To ensure data quality, CrODA-gator uses RTDP. Unlike simple boundary checks, RTDP compares a new sensor entry against:
- Pre-defined Min/Max thresholds.
- Local Consensus: The average value of neighboring data points within a specific radius.
- Reputation: Devices that consistently submit outliers are eventually banned via their Unique Device Identifier (UDID).
Experimental Results
The platform was tested using datasets ranging from 3,000 to 50,000 entries, focusing on WiFi RSSI and atmospheric pressure data.
- Mapping Efficiency: The RIE algorithm showed remarkable scalability. For large datasets, it executed in approximately 186-284ms, whereas the baseline Basic Recursive Indexing (BRI) took nearly 1 second.
- Filter Robustness: While RTDP is slightly slower than basic filtering (due to the neighborhood average calculation), the execution time increase was 31x relative to a 16.7x data size increase, outperforming the baseline which increased by 38x. This suggests RTDP scales more efficiently under heavy loads.

Critical Insight & Conclusion
The true value of CrODA-gator lies not just in its algorithms, but in its Open Access philosophy. By offering well-documented APIs and bulk import features, it lowers the barrier to entry for educational and corporate entities.
Limitations: Currently, the platform relies on UDID for device identification, which may raise privacy concerns despite non-sensitive data collection. Future iterations would benefit from decentralized identity or differential privacy techniques.
Future Outlook: As the Internet of Things (IoT) expands, platforms like CrODA-gator will transition from "Mobile Crowdsensing" to "Universal Sensing," integrating data from wearables and smart home devices to create a high-fidelity digital twin of our environment.
