ADDCM: Solving Data Scarcity in Industrial IoT via Adaptive Dropout and Crowdsourcing
9447_An Adaptive Dropout Deep Computation Model for Industrial IoT Big Data Learning With Crowdsourcing to Cloud Computing.
The paper proposes the Adaptive Dropout Deep Computation Model (ADDCM), a tensor-based deep learning framework integrated with cloud-based crowdsourcing for Industrial IoT. It achieves state-of-the-art feature learning performance by dynamically adjusting dropout rates and utilizing human intelligence to solve the labeled data scarcity problem.
TL;DR
In the world of Industrial IoT (IIoT), we have a "Data-Label Paradox": massive amounts of raw sensor data but a crippling shortage of labeled training samples. This paper introduces the Adaptive Dropout Deep Computation Model (ADDCM). It utilizes a tensor-based architecture to handle big data, an adaptive dropout mechanism to kill overfitting, and a cloud-based crowdsourcing pipeline that uses "Human Intelligence" to refine its parameters.
Background & Motivation: The Overfitting Trap
Industrial IoT systems collect heterogeneous big data (images, audio, vibration signals) that are naturally represented as high-order tensors. While Deep Computation Models (DCM) generalize traditional neural networks into the tensor space, they are notoriously prone to overfitting.
When labeled data is scarce, the model "memorizes" the noise. Standard solutions like Dropout help, but they treat every layer the same—usually with a flat 50% keep rate. The authors argue that this ignores the inductive bias of deep networks: lower layers learn general features, while higher layers learn specific tasks.
Methodology: The Three Pillars of ADDCM
1. Adaptive Dropout Distribution
Instead of a fixed rate, the authors propose a distribution function based on the normal distribution shifted for layer position :
This creates a monotonically decreasing activation rate. Lower layers (closer to input) stay more active to capture raw features, while higher layers become sparser to prevent co-adaptation of complex feature detectors.
Figure 1: The framework of ADDCM with cloud-based crowdsourcing.
2. Maximum Entropy Outsourcing
Not all unlabeled data is worth paying humans to label. The model uses Uncertainty Sampling via Maximum Entropy:
Samples with the highest entropy (where the model is most confused) are outsourced to the cloud for human labeling, maximizing the "information gain" per dollar spent.
3. Extended SLME for Multi-labeling
Crowdsourced labels are noisy. To handle this, the authors extended the Supervised Learning from Multiple Experts (SLME) algorithm. By introducing a confusion matrix for each worker, the system iteratively estimates both the ground truth and the reliability (expertise) of each worker using an EM (Expectation-Maximization) algorithm.
Experimental Validation
The model was tested on CUAVE (multimodal digits) and SNAE2 (industrial video) datasets.
Key Findings:
- Overfitting Prevention: As layers increased from 3 to 5, standard DCM performance plummeted (error rose from 19.6% to 31.3%). ADDCM remained stable, proving the effectiveness of the adaptive dropout.
- Crowdsourcing Efficiency: The ADDCM-DDC (Crowdsourced) model achieved a Top-2 error of 9.0%, nearly matching the Supervised version (ADDCM-SU) at 7.5%, and significantly outperforming the unsupervised baseline of 11.5%.
Table 1: Top-1 Error rates comparing DCM, DDCM, and ADDCM.
Critical Insight: Why it Works
The brilliance of this work lies in the synergy between math and crowd. The adaptive dropout provides a structural regularization that respects the hierarchical nature of feature learning. Meanwhile, the Maximum Entropy selection ensures that the expensive human intelligence is only used where the "structural math" fails to provide certainty. It's a textbook example of Active Learning applied to high-dimensional tensor spaces.
Conclusion & Future Look
ADDCM represents a significant step for Industrial IoT analytics where "Big Data" is often "Thin Data" (lots of volume, few labels). Future work could improve this by optimizing the initial parameter sets using Meta-Learning or Self-Supervised pre-training (like Masked Autoencoders) to further reduce the reliance on crowdsourcing.
Takeaways
- Adaptive Dropout > Static Dropout for deep structures.
- Tensor Decomposition is essential for Industrial Big Data.
- Human-in-the-loop is not a sign of failure, but a strategic tool for SOTA performance.
