Engendering Trust: Can AI Feedback Outperform Human Experts in Gesture Learning?
Engendering Trust in Automated Feedback: A Two Step Comparison of Feedbacks in Gesture Based Learning
The paper presents a comparative study of automated vs. manual feedback in American Sign Language (ASL) learning using the ASLHelp application. It proposes a transition from generic "correct/incorrect" labels to explainable, concept-level feedback based on location, movement, and handshape, achieving an agreement rate with human experts in nearly half of the cases.
Executive Summary
TL;DR: This research tackles the "trust gap" in AI education by comparing how a fine-grained, automated feedback system (ASLHelp) stacks up against human experts in American Sign Language (ASL) instruction. By breaking gestures down into specific linguistic concepts—Location, Handshape, and Movement—the authors demonstrate that AI feedback is not only comparable to human experts but is often more precise in identifying subtle execution errors.
Background: Positioned at the intersection of Computer-Aided Language Learning (CALL) and Explainable AI (XAI), this work moves beyond simple "right/wrong" binary classification. It establishes a framework for trust by making the "reasoning" of the AI transparent to the learner.
The Problem: The "Black Box" of Automated Feedback
In traditional gesture-based learning (like ASL, surgery, or sports), manual feedback from a human teacher remains the "gold standard." Learners often view AI feedback with apprehension because:
- Lack of Transparency: They don't know how the machine decided their gesture was wrong.
- Perceived Inaccuracy: The belief that a machine cannot capture the "nuance" of human movement.
Most prior work in gesture recognition focuses on accuracy of identification, but fails to provide formative assessment—the "why" and "how" that helps a student improve.
Methodology: Concept-Level Deconstruction
The core innovation lies in the Grammar Expression of Gesture. Instead of treating a sign as a single data point, the authors define a gesture (GE) using a context-free grammar:
System Architecture
The ASLHelp system uses PoseNet for keypoint estimation (tracking wrists, shoulders, and elbows) and Dynamic Time Warping (DTW) to handle variations in speed between a novice and an expert.

- Location Recognition: Uses a "bucket" system to track wrist joints in 2D space.
- Movement: Employs DTW and Z-normalization to compare trajectory regardless of the learner's distance from the camera.
- Handshape: Uses a specialized CNN to analyze cropped images of the hands.
Experiments: The Two-Step Blind Test
To measure trust and effectiveness, the authors designed a rigorous two-step evaluation:
- Direct Comparison: Comparing AI feedback against manual feedback from three experts across 154 videos.
- Blind Appropriateness Test: A fourth "second-level" expert, unaware of the feedback source, chose which feedback (AI or Manual) was more appropriate for a specific video.
Key Results
The findings challenge the assumption that human feedback is always superior:
- High Alignment: In 78.87% of cases, the AI and humans agreed on at least 2 out of 3 components.
- The Precision Gap: Second-level experts chose AI feedback 40.91% of the time.
- Sensitivity: Humans were more "forgiving," often ignoring minor mistakes, whereas the AI was precise—a trait valuable for professional or technical training.

Critical Insight & Analysis
The disparity in results often came down to Handshape. Automated systems struggled more with handshapes due to lighting and background noise in student-recorded videos. However, the study proves that when the AI does provide feedback, it is often more granular than a human's "gut feeling."
Takeaways for the Industry:
- XAI is non-negotiable: Users only trust AI when it uses the same "rubrics" as humans (e.g., specific concepts like Location/Movement).
- Hybrid Utility: AI is currently best suited as a "precision tool" that complements human instruction by catching errors a teacher might overlook during a fast-paced session.
Conclusion
This work demonstrates that engendering trust in AI isn't just about higher accuracy percentages; it's about alignment in communication. By adopting the linguistic structures of ASL (Stokoe's modalities), ASLHelp provides a roadmap for how AI can move from a simple critic to a sophisticated, explainable mentor.
Future work aims to calibrate the system for varying gesture difficulty levels and improve handshape recognition in unconstrained environmental conditions.
