Flash the Fish: High-Quality Video Annotation Through Gamified Crowdsourcing
A crowdsourcing approach to support video annotation
This paper introduces "Flash the Fish," an innovative crowdsourcing framework that converts video annotation into an engaging online game to generate large-scale ground truth data. By processing noisy player clicks and feeding them as seeds into Region Growing and GrabCut segmentation algorithms, the system produces high-quality object contours for computer vision tasks.
TL;DR
Generating ground truth data for computer vision is a notoriously "boring" task that consumes countless research hours. This paper proposes a solution by turning video annotation into a fun, underwater-themed Flash game. By intelligently processing "noisy" player clicks and feeding them into classic segmentation algorithms like GrabCut and Region Growing, the authors successfully transformed casual gaming into professional-grade object segmentation data with an F1-measure of up to 80%.
Problem & Motivation: The Annotation Bottleneck
In the world of computer vision, an algorithm is only as good as the ground truth (GT) used to evaluate it. However, creating these annotations — especially frame-by-frame video contours — is a soul-crushing task.
Existing solutions fall into two unsatisfying camps:
- The Professional Route: Tools like ViPER-GT or LabelMe require dedicated researchers to draw polygons manually. It’s accurate but impossible to scale.
- The Crowdsourcing Route: Platforms like Amazon Mechanical Turk offer scale, but quality is often abysmal because users lack motivation or expertise.
The authors' insight was simple: What if we hide the work inside a game? By gamifying the process, they can attract a massive number of users. The challenge then shifts from "how to get data" to "how to extract signal from the noise of thousands of uncoordinated clicks."
Methodology: From Clicks to Contours
The proposed pipeline consists of three main stages: Data Collection, Noise Filtering, and Algorithmic Refinement.
1. The Game: Flash the Fish
Users play a game where they must "click" on moving fish to take photos and score points. As they progress through levels, the frame rate increases and the time limit decreases, naturally separating casual players from "expert" annotators.

2. Processing the Noise
Raw clicks are just (x, y) coordinates. To make sense of them, the authors use:
- K-Means Clustering: To identify high-density areas that likely represent actual fish.
- Heatmaps: Generated by summing 3D Gaussian distributions for each click, helping to identify the most probable center of an object.
3. Segmentation Refinement
The paper compares two strategies to turn points into shapes:
- Region Growing: Uses the heatmap peaks as "seeds." It works by expanding from a pixel to its neighbors based on color similarity.
- GrabCut: This is more robust. The system draws a Convex Hull (a boundary) around a cluster of clicks. This area is fed into the GrabCut algorithm, which uses a probabilistic graph-cut approach to find the exact object boundary.

Experiments & Results: Does Gaming Really Work?
The authors ran a 4-day Facebook event, resulting in over 260,000 clicks from 80 users. They compared the machine-generated segments against a hand-labeled dataset of 4,140 objects.
| Level | Acquired Clicks | GrabCut F1-Score | Region Growing F1-Score |
|---|---|---|---|
| 1 (Easy) | 71,105 | 0.79 | 0.23 |
| 7 (Hard) | 342 | 0.23 | 0.66 |
Key Insights:
- GrabCut is the "Volume King": In early levels with massive amounts of clicks, GrabCut excelled at ignoring the background and finding the fish, achieving high Precision and Recall.
- Region Growing is the "Expert Specialist": In high-difficulty levels, only the best players survived. While there were fewer clicks, they were extremely precise. Region Growing performed better here because the "seed" was almost always perfectly centered on the object.

Critical Analysis & Conclusion
The value of this research lies in its data-fusion strategy. It acknowledges that crowdsourced data is fundamentally "noisy" and uses established computer vision algorithms (GrabCut) to fix human imprecision.
Limitations:
- Temporal Consistency: The current method processes frames individually. It does not yet fully exploit the temporal link between frames in a video (tracking).
- Algorithm Sensitivity: As seen in the results, GrabCut's performance collapses if the number of clicks is too low.
Future Outlook:
This approach is particularly powerful for "Big Data" domains like environmental monitoring (underwater videos) or traffic analysis, where the sheer volume of video makes manual annotation impossible. By aligning human entertainment with scientific needs, we can create self-sustaining annotation ecosystems.
