Intelligent Data Augmentation: Streamlining VQA Dataset Expansion via Crowdsourcing

A Crowdsourcing Tool for Data Augmentation in Visual estion Answering Tasks

Ramon Silva, Augusto Fonseca, Ronaldo Goldschmidt
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a specialized crowdsourcing tool designed for data augmentation in Visual Question Answering (VQA) tasks, specifically focusing on binary (yes/no) questions. By integrating Siamese Networks, SSD object detection, and Word2Vec, the tool intelligently filters candidate images and questions to optimize the human curation process.

TL;DR

The paper presents a collaborative crowdsourcing framework designed to tackle the "data hunger" of Visual Question Answering (VQA) models. By utilizing deep learning modules (Siamese Networks, SSD, and Word2Vec) to pre-filter and rank image-question pairs, the tool reduces the curation workload by over 1,000 times compared to brute-force labeling, specifically for binary VQA tasks.

Background and Motivation

Visual Question Answering is the "holy grail" of multi-modal AI, requiring a model to understand both visual scenes and natural language nuances. However, the state-of-the-art—Deep Neural Networks—suffers from a harsh reality: gaining a 10-12% boost in accuracy often requires a 10x increase in dataset size.

Manual labeling at this scale is prohibitively expensive. While "blind" data augmentation (like flipping images) exists, it often leads to redundant data that offers diminishing returns. The authors identified a crucial need for a "disciplined augmentation" approach—one that finds new, relevant images from external sources (like ImageNet) and matches them with appropriate questions to be verified by humans.

Methodology: The Four-Stage Pipeline

The system's architecture is built on a series of intelligent filters designed to maximize the "value per click" for human curators.

1. Image Filtering (The Siamese Gate)

To ensure the new images () are relevant to the original task (), the tool uses a Siamese Neural Network. This network computes a similarity score () between pairs of images. Only images that cross a similarity threshold () move forward.

2. Semantic Extraction

The tool then performs a dual-modality analysis:

  • Visual Side: Uses Single Shot Detection (SSD) to identify objects within the new images.
  • Textual Side: Uses a POS Tagger to extract nouns from existing questions in the original dataset.

3. Question Filtering (Word Embeddings)

Using Word2Vec, the tool measures the semantic distance () between the objects identified in the image and the nouns in the question. If a question is semantically "too far" from the image content, it is discarded.

4. Prioritized Curation

The remaining pairs are presented to humans, ranked by their relevance. The curator simply confirms if a question applies to the image and provides the final binary answer.

System Architecture Figure 1: The architecture of the curation tool, showcasing the modular flow from raw data to human verification.

Experimental Results

In a practical instantiation targeting "Dog" and "Cat" categories from ImageNet and the VQA dataset, the numbers were striking:

  • Initial Candidate Pairs: ~341 Million
  • Post-Filtering Pairs: 259,716
  • Efficiency Gain: ~1,300x reduction in human effort.
  • Human Performance: Curators took about 12 seconds per item, suggesting a highly focused and low-friction interface.

SSD Object Detection Example Figure 2: An example of the SSD model detecting objects (Bike, Car, Dog) to generate terms for semantic matching.

Critical Insight & Future Outlook

The brilliance of this work lies in its Search-Space Reduction. Instead of asking humans to "write a question for this image" (high cognitive load), it asks them to "verify if this existing question fits this similar image" (low cognitive load).

Limitations & Future Work

  • Binary Focus: The current tool is limited to "Yes/No" questions. Expanding to open-ended questions remains a challenge due to the increased complexity of answer verification.
  • Active Learning: The authors plan to incorporate Active Learning to prioritize not just "similar" images, but those that the VQA model is currently "uncertain" about.

Conclusion

This tool represents a significant step toward scalable VQA dataset creation. By treating data augmentation as an intelligent filtering problem rather than a generative one, it paves the way for high-quality, large-scale multimodal datasets that are both semantically diverse and labor-efficient.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize active learning strategies to further minimize human labeling effort in Visual Question Answering datasets.
  • Which paper originally proposed the Siamese Network architecture for image similarity, and how has its application in data curation evolved since 2018?
  • Explore how modern Large Language Models (LLMs) and Multimodal Models (LVMs) are currently being used to automate VQA data augmentation compared to the crowdsourcing approach described here.
Contents
Intelligent Data Augmentation: Streamlining VQA Dataset Expansion via Crowdsourcing
1. TL;DR
2. Background and Motivation
3. Methodology: The Four-Stage Pipeline
3.1. 1. Image Filtering (The Siamese Gate)
3.2. 2. Semantic Extraction
3.3. 3. Question Filtering (Word Embeddings)
3.4. 4. Prioritized Curation
4. Experimental Results
5. Critical Insight & Future Outlook
5.1. Limitations & Future Work
6. Conclusion