DPDCM: Capturing Multimodal Interactions with Double-Projection and Privacy-Preserving Cloud Computing
10387_Privacy-Preserving Double-Projection Deep Computation Model With Crowdsourcing on Cloud for Big Data Feature Learning.
This paper introduces the Double-Projection Deep Computation Model (DPDCM) and its privacy-preserving variant (PPDPDCM) for big data feature learning in IoT. By projecting input data into dual separate nonlinear subspaces to capture interactive multimodal correlations, the model achieves state-of-the-art classification accuracy on Animal-20 and NUS-WIDE-14 datasets compared to conventional Deep Computation Models (DCM).
TL;DR
The Double-Projection Deep Computation Model (DPDCM) addresses the challenge of heterogeneous Big Data feature learning by replacing standard hidden layers with dual-subspace projections. By leveraging the Kronecker product to model interactions between modalities and employing BGV homomorphic encryption for secure cloud-based training, the authors achieve a 4% accuracy boost and a 40% reduction in training time compared to traditional tensor-based models.
Problem & Motivation: Beyond Simple Feature Concatenation
In the era of the Internet of Things (IoT), data is inherently multimodal (images, text, video) and multidomain. Existing models like Multimodal Deep Boltzmann Machines often learn features for each modality separately and then simply concatenate them.
The Critical Flaw: Concatenation is too simplistic to reveal the deep "interactive correlations" between modalities. For instance, in a game video, the visual stream and the annotation text are not just parallel data points; they represent the same event through different nonlinear transformations. Standard Deep Computation Models (DCM) fail to capture these subspace-level interactions, leading to suboptimal feature representation.
Methodology: The Power of Double Projection
The heart of this paper is the Double-Projection Layer. Instead of mapping an input tensor to a single hidden state , the DPDCM maps it into two separate hidden subspaces, and .
1. Dual-Subspace Architecture
The model uses an interaction representation defined by the Kronecker product of the two subspaces: This allows the model to assign weights to every possible combination of features from the two subspaces, effectively capturing the nonlinear "dialogue" between different data aspects.
Figure 1: Comparison between (a) Conventional Tensor Auto-encoder and (b/c) the proposed Double-Projection Tensor Auto-encoder.
2. Privacy-Preserving Cloud Training (PPDPDCM)
To handle Big Data volumes, the authors outsource computation to the cloud. To prevent data leakage, they utilize the BGV (Brakerski-Gentry-Vaikuntanathan) scheme. Since BGV does not natively support the division or exponentiation required for the Sigmoid activation function, the authors use a polynomial approximation: This allows the entire back-propagation process to occur on encrypted ciphertexts.
Experiments & Results
The model was evaluated on the Animal-20 and NUS-WIDE-14 datasets, comparing a standard DCM against the DPDCM with varying numbers of hidden layers.
Performance Gains
- Accuracy: On NUS-WIDE-14, DPDCM achieved 77.2% accuracy compared to 73.4% for DCM.
- Optimal Depth: Both models performed best with 3 hidden layers, suggesting a "sweet spot" for hierarchical feature abstraction in these datasets.
Table 1: Classification results on NUS-WIDE-14 across different initializations.
Efficiency vs. Privacy
The Privacy-Preserving (PPDPDCM) variant showed a massive improvement in training efficiency. By utilizing cloud resources, the training time was slashed by approximately 40%, with only a ~1% drop in accuracy—a negligible trade-off for the security and speed gained.
Table 2: Comparison of training time (seconds) between local DPDCM and cloud-crowdsourced PPDPDCM.
Critical Insight & Conclusion
The DPDCM succeeds because it treats multimodal data not as a collection of vectors, but as interacting manifolds. The use of high-order tensor spaces preserves the structural integrity of Big Data that vector-based models often destroy.
Limitations: While the sigmoid approximation works well, more complex activation functions (like ReLU or Swish) might be harder to implement under BGV without significant overhead. Furthermore, the 1% accuracy loss, while small, could be significant in high-stakes fields like medical diagnostics.
Final Takeaway: This research provides a robust framework for scalable, secure, and interaction-aware deep learning, marking a significant step forward for Industrial IoT and privacy-sensitive Big Data analytics.
