CrowdBuy: Unlocking Private Mobile Data for High-Quality Machine Learning
CrowdBuy: Privacy-friendly Image Dataset Purchasing via Crowdsourcing
The paper introduces CrowdBuy and CrowdBuy++, a crowdsourcing-based framework for purchasing high-quality image datasets while preserving user privacy and data ownership. It utilizes CNN-extracted features and an autoencoder to perform privacy-preserving image selection, achieving State-of-the-Art (SOTA) efficiency and utility in data acquisition for deep learning tasks.
TL;DR
CrowdBuy is a novel crowdsourcing framework that enables buyers to purchase high-quality image datasets directly from mobile users' private albums. By utilizing deep learning features instead of raw images for the selection process, it ensures image ownership, privacy, and truthful pricing, while maximizing dataset diversity for better model generalization.
Context: The Data Bottleneck
In the era of Deep Learning, the hunger for high-quality, labeled data is insatiable. While web scraping is the current standard, it suffers from legal watermarking, privacy violations, and a lack of diversity in rare categories (e.g., specific medical conditions or biometric data). Ironically, billions of high-quality images sit idle on mobile devices. Why haven't we used them? Because users fear privacy leakage, and buyers fear low-quality, fraudulent data.
The Core Insight: Selection in Feature Space
CrowdBuy bridges this gap by moving the "browsing" phase from the raw image domain to a compressed 2D feature space.
- Feature Extraction: Sellers use a pre-trained VGG-16 model to extract
fc8layer features (1000 dimensions) and then use an Autoencoder to compress this to a 2D vector. - Privacy-Preserving Bidding: Instead of sending the photo, the seller sends the 2D vector and a price.
- Selection Logic: The cloud server identifies "matched" images by measuring the Euclidean distance between the seller's vector and the buyer's sample image vector.

Methodology: Beyond Simple Matching
CrowdBuy doesn't just look for "similar" images. It introduces three quality metrics to ensure the buyer gets their money's worth:
- Quantity: Maximum images for a fixed budget.
- Matching Degree: Highest similarity to the target (e.g., different angles of the same dog breed).
- Diversity (The Secret Sauce): Using a hypercube coverage model in the 2D space, the system selects images that are spread out. This prevents the "redundancy" problem where a buyer pays for 1,000 nearly identical photos.
CrowdBuy++ and Feature Indistinguishability
For extreme privacy (Level-2), the authors propose CrowdBuy++. Even 2D features can sometimes be inverted to reveal silhouettes. CrowdBuy++ uses a MinHash based strategy to achieve feature-indistinguishability—a generalization of Differential Privacy. It masks the exact feature vector location while allowing the cloud to still calculate overlaps and set coverage.
Experimental Validation: Quality & Efficiency
Theoretical privacy is useless if the resulting dataset is junk. The researchers tested CrowdBuy on 222,300 images from ImageNet.
1. Robustness to Malicious Users
Can a user upload a picture of a "cat" when the buyer asked for a "dog"? CrowdBuy's distance threshold (R) ensures that even with 50% malicious noise, the precision of the final dataset remains above 90%.

2. Generalization Gains
The most striking result is the impact of the Diversity Metric. A CNN trained on a "Diverse" dataset selected by CrowdBuy reached 99.2% accuracy on unknown categories, whereas a "Similarity-only" dataset plummeted to roughly 50%. This proves that selecting "different" matched images is far more valuable than selecting "closest" matched images.

Critical Insight & Future Outlook
CrowdBuy succeeds because it treats data as a commodity with a budget. By implementing truthful incentive mechanisms (where sellers are best off reporting their real costs), it solves the economic game-theory problem of crowdsourcing.
Limitations: The current system relies on a pre-trained VGG-16 and Autoencoder. If the buyer's requested category is too far outside the distribution of these pre-trained models, the feature extraction might lose "semantic resolution."
Future Work: We expect to see this framework extended to Multimodal data (Audio/Video) and perhaps integrated with Blockchain for immutable proof of ownership and automated payments via smart contracts.
Summary Table
| Feature | CrowdBuy | CrowdBuy++ |
|---|---|---|
| Privacy Level | Level-1 (Features exposed) | Level-2 (Features masked) |
| Quality Logic | Submodular Optimization | MinHash Intersection |
| Truthfulness | Guaranteed (Myerson Lemma) | Guaranteed |
| Inference Time | ~1.4s per image | ~1.5s per image |
