CrowdBuy: Unlocking Private Mobile Data for High-Quality Machine Learning

CrowdBuy: Privacy-friendly Image Dataset Purchasing via Crowdsourcing

2018-04-01
Lan Zhang, Yannan Li, Xiang Xiao, Xiang-Yang Li, Junjun Wang, Anxin Zhou, Qiang Li
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces CrowdBuy and CrowdBuy++, a crowdsourcing-based framework for purchasing high-quality image datasets while preserving user privacy and data ownership. It utilizes CNN-extracted features and an autoencoder to perform privacy-preserving image selection, achieving State-of-the-Art (SOTA) efficiency and utility in data acquisition for deep learning tasks.

TL;DR

CrowdBuy is a novel crowdsourcing framework that enables buyers to purchase high-quality image datasets directly from mobile users' private albums. By utilizing deep learning features instead of raw images for the selection process, it ensures image ownership, privacy, and truthful pricing, while maximizing dataset diversity for better model generalization.

Context: The Data Bottleneck

In the era of Deep Learning, the hunger for high-quality, labeled data is insatiable. While web scraping is the current standard, it suffers from legal watermarking, privacy violations, and a lack of diversity in rare categories (e.g., specific medical conditions or biometric data). Ironically, billions of high-quality images sit idle on mobile devices. Why haven't we used them? Because users fear privacy leakage, and buyers fear low-quality, fraudulent data.

The Core Insight: Selection in Feature Space

CrowdBuy bridges this gap by moving the "browsing" phase from the raw image domain to a compressed 2D feature space.

  1. Feature Extraction: Sellers use a pre-trained VGG-16 model to extract fc8 layer features (1000 dimensions) and then use an Autoencoder to compress this to a 2D vector.
  2. Privacy-Preserving Bidding: Instead of sending the photo, the seller sends the 2D vector and a price.
  3. Selection Logic: The cloud server identifies "matched" images by measuring the Euclidean distance between the seller's vector and the buyer's sample image vector.

CrowdBuy Overall Architecture

Methodology: Beyond Simple Matching

CrowdBuy doesn't just look for "similar" images. It introduces three quality metrics to ensure the buyer gets their money's worth:

  • Quantity: Maximum images for a fixed budget.
  • Matching Degree: Highest similarity to the target (e.g., different angles of the same dog breed).
  • Diversity (The Secret Sauce): Using a hypercube coverage model in the 2D space, the system selects images that are spread out. This prevents the "redundancy" problem where a buyer pays for 1,000 nearly identical photos.

CrowdBuy++ and Feature Indistinguishability

For extreme privacy (Level-2), the authors propose CrowdBuy++. Even 2D features can sometimes be inverted to reveal silhouettes. CrowdBuy++ uses a MinHash based strategy to achieve feature-indistinguishability—a generalization of Differential Privacy. It masks the exact feature vector location while allowing the cloud to still calculate overlaps and set coverage.

Experimental Validation: Quality & Efficiency

Theoretical privacy is useless if the resulting dataset is junk. The researchers tested CrowdBuy on 222,300 images from ImageNet.

1. Robustness to Malicious Users

Can a user upload a picture of a "cat" when the buyer asked for a "dog"? CrowdBuy's distance threshold (R) ensures that even with 50% malicious noise, the precision of the final dataset remains above 90%.

Precision against Noise

2. Generalization Gains

The most striking result is the impact of the Diversity Metric. A CNN trained on a "Diverse" dataset selected by CrowdBuy reached 99.2% accuracy on unknown categories, whereas a "Similarity-only" dataset plummeted to roughly 50%. This proves that selecting "different" matched images is far more valuable than selecting "closest" matched images.

Accuracy Comparisons

Critical Insight & Future Outlook

CrowdBuy succeeds because it treats data as a commodity with a budget. By implementing truthful incentive mechanisms (where sellers are best off reporting their real costs), it solves the economic game-theory problem of crowdsourcing.

Limitations: The current system relies on a pre-trained VGG-16 and Autoencoder. If the buyer's requested category is too far outside the distribution of these pre-trained models, the feature extraction might lose "semantic resolution."

Future Work: We expect to see this framework extended to Multimodal data (Audio/Video) and perhaps integrated with Blockchain for immutable proof of ownership and automated payments via smart contracts.

Summary Table

FeatureCrowdBuyCrowdBuy++
Privacy LevelLevel-1 (Features exposed)Level-2 (Features masked)
Quality LogicSubmodular OptimizationMinHash Intersection
TruthfulnessGuaranteed (Myerson Lemma)Guaranteed
Inference Time~1.4s per image~1.5s per image

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Differential Privacy or Local Differential Privacy in the context of image feature protection for machine learning data collection.
  • Which original research established the BEACON mechanism for budget-feasible incentive design, and how does CrowdBuy adapt it for non-numerical image utilities?
  • Explore if MinHash-based feature indistinguishability has been applied to other modalities such as audio or biometric data crowdsourcing.
Contents
CrowdBuy: Unlocking Private Mobile Data for High-Quality Machine Learning
1. TL;DR
2. Context: The Data Bottleneck
3. The Core Insight: Selection in Feature Space
4. Methodology: Beyond Simple Matching
4.1. CrowdBuy++ and Feature Indistinguishability
5. Experimental Validation: Quality & Efficiency
5.1. 1. Robustness to Malicious Users
5.2. 2. Generalization Gains
6. Critical Insight & Future Outlook
7. Summary Table