Crowd Tasker: Breaking the Screen Barrier in Human Computation
"Hi! I am the Crowd Tasker" Crowdsourcing through Digital Voice Assistants
This paper introduces "Crowd Tasker," the first system enabling crowdsourcing through Digital Voice Assistants (DVAs) like Google Assistant. It validates the feasibility of voice-only crowd work across six task types, demonstrating that native English speakers achieve performance comparable to traditional web interfaces while gaining significant flexibility.
TL;DR
Is it possible to "work" while folding laundry or cooking dinner? Researchers from the University of Melbourne and UCL have introduced Crowd Tasker, a system that transforms Digital Voice Assistants (DVAs) into crowdsourcing hubs. Their findings suggest that for many tasks, your voice is just as accurate as your keyboard, but significantly more flexible.
The Problem: The High Friction of Micro-tasks
Despite the "micro" nature of modern crowd work, the friction to start is surprisingly high. Workers must sit at a desk, log in, browse lists, and be visually tethered to a screen. This "interaction tax" prevents people from utilizing small pockets of time—what scholars call Cognitive Surplus.
Previous attempts at voice-based work were limited to basic phone-in systems or specialized transcription apps. There was no unified, hands-free way to browse, launch, and execute a variety of tasks via the smart speakers already sitting in millions of homes.
Methodology: Designing for the Ear, Not the Eye
The researchers built "Crowd Tasker" on Google Assistant. The core challenge was translating visual workflows into a logical conversational flow. They tested two categories:
- Voice-Compatible: Traditionally text-based tasks (Sentiment Analysis, Text Moderation).
- Voice-Based: Tasks involving audio stimuli (Audio Annotation, Transcription).
The system architecture utilized a sequence of Intents—mapping specific vocal commands to actions like "Start a Task" or "Check Progress."
Figure 1: The session flow and intent mapping that allows hands-free navigation of tasks.
Experiments: Lab Control vs. Living Room Chaos
The researchers conducted a two-phase evaluation:
- Lab Study (N=30): A direct "A/B test" comparing a web UI to a Google Home speaker.
- Field Deployment (N=12): A week-long study where participants used the device in their actual homes.
Key Insights from the Data:
- The "Native" Advantage: While native English speakers had parity between web and voice, non-native speakers struggled due to both the DVA's recognition errors and the synthetic voice's difficulty to parse.
- Speed Efficiency: For tasks like Speech Transcription, voice was actually faster than typing.
Figure 2: Comparison of worker accuracy and task time. While accuracy is slightly lower in complex tasks (like Comprehension), the speed gains in transcription are evident.
Designing for Voice: The "Golden Rules"
The study concludes with a set of crucial design guidelines for future "Voice-First" work systems:
- Question First, Stimulus Second: In voice, users need to know what to listen for before the audio clip plays to reduce working memory load.
- Minimize Options: Tasks with more than 3-4 options (like 6-way emotion labeling) lead to higher error rates. Keep it binary (Spam/Not Spam) where possible.
- The "Repeat" Control: Users need granular control to repeat specific parts of a task rather than hearing the whole prompt again.
Critical Analysis & Future Outlook
Crowd Tasker proves that the smart speaker is no longer just for setting timers; it is a viable "office" for micro-contributions.
Limitations: The reliance on Google Assistant's proprietary speech-to-text means that users with accents remain marginalized—a significant "Inductive Bias" in current DVA systems.
Future Impact: As companies like Amazon (who own both Alexa and MTurk) look for ways to expand their workforce, voice interfaces offer a path to include workers who are multitaskers, commuters, or those with visual/motor disabilities. The future of crowd work isn't just "on the go"—it's "on the air."
Takeaway
Voice-based crowdsourcing is best for short, low-complexity tasks with limited answer options. It succeeds not by replacing the web, but by capturing the "idle moments" of our lives.
