Content Is King: Crafting Crowdsourced Tasks for Human-Robot Interaction
Content Is King: Impact of Task Design for Eliciting Participant Agreement in Crowdsourcing for HRI
This study investigates how task design in crowdsourcing influences participant agreement when interpreting non-anthropomorphic drone flight paths. By comparing various question types and task lengths on Amazon MTurk, the researchers identified best practices to elicit consistent "action-based" responses for Human-Robot Interaction (HRI).
TL;DR
When designing crowdsourced studies for drones or simple robots, the visual content (what the robot does) matters far more for participant agreement than the specific phrasing of your questions. This study reveals that participants quickly habituate to task structures, but asking multiple complementary questions doesn't necessarily degrade data quality—provided you vary the visual length of your prompts to keep them engaged.
The "Vibe" Over the "Verbiage": Why This Matters
In the world of Human-Robot Interaction (HRI), we often try to make non-humanoid robots (like drones) "talk" through movement. But how do we know if people actually understand what a drone is trying to say? Crowdsourcing via platforms like Amazon MTurk is a great way to get diverse data, but it comes with a massive headache: Rater Agreement. If 100 people watch a drone spiral and give 100 different interpretations, the data is useless.
The authors of this paper set out to find the "Goldilocks zone" of task design. They wanted to know: Does changing a question from "What would this drone say?" to "What gesture is this?" actually change the result? And does making the study shorter actually result in better data?
Methodology: Testing the Parameters of Perception
The researchers tested 80 "mTurk Masters" using 16 unique drone flight videos (e.g., "X-shape," "Horizontal Figure 8," "Hover").
The Variables:
- Question Types: Speech-based, Gesture-based, and Physical response-based.
- Task Length: Including or excluding the PANAS (Positive and Negative Affect Scale).
- Repetition: Asking one vs. two questions per video.

Core Insights: What the Data Tells Us
1. The "Visual Habituation" Trap
One of the most striking findings was regarding attention checks. Participants were significantly more likely to fail an attention check if the question looked visually similar (in length and shape) to previous questions.
- The Insight: Participants aren't necessarily lazy; they are efficient. If every page has a one-line question of ~50 characters, their brains "autopilot" through the text. The researchers recommend varying question lengths to force the brain to re-engage with the prompt.
2. Doubling the Questions, Not the Pain
Common wisdom suggests that more questions = more fatigue = worse data. However, this study found:
- Doubling the questions only added about 7 minutes to a 30-minute task.
- Adding two questions about the same video often provided complementary information (e.g., one question elicited "the drone is telling me to move," while the other confirmed "it sounds like a command").
- Reducing the pre-survey questionnaires only saved 2.4 minutes—hardly enough to justify losing valid demographic or psychological data.
3. Content is the Real Driver
When the researchers looked at Cohen’s Kappa (a measure of agreement), the agreement scores for specific videos were consistently "Substantial" or "Near Perfect." In contrast, agreement scores for questions were highly variable.

- The Interpretation: If a drone's movement is clear (like a "Back and Forth" motion), people will agree on its meaning regardless of whether you ask them what the drone "says" or what "gesture" it is making. The Flight Path (Content) is the primary determinant of communication success.
Critical Analysis & Professional Recommendations
Takeaways for HRI Researchers:
- Stop Over-Optimizing the Questionnaire: If your rater agreement is low, the problem is likely your robot's motion primitive, not your survey's wording.
- Triangulate Your Data: Don't be afraid to ask two or three questions about the same stimulus. Participants seem willing to provide diverse insights for the same video.
- Engagement via Variation: Change the visual layout and character length of your questions to prevent MTurkers from falling into a "zombie" state of clicking through.
Limitations:
The study focused on "Action-Based" responses (e.g., "The drone wants me to move"). It did not dive deep into the emotional nuances of the responses, which is a key area for future work (e.g., does a specific flight path make a user feel "threatened" versus "guided"?).
Conclusion
This work provides a vital scaffold for anyone designing remote HRI experiments. By prioritizing clear robotic "gestures" and maintaining engagement through visual prompt variation, researchers can achieve high-quality, actionable data without sacrificing the depth of their surveys.
