[CHI/ASSETS] VRSL: Breaking the Silence in VR with 360-Degree Sign Language Feeds

VRSL:Exploring the Comprehensibility of 360-Degree Camera Feeds for Sign Language Communication in Virtual Reality

Summary
Problem
Method
Results
Takeaways
Abstract

VRSL explores the use of body-mounted 360-degree camera feeds to enable American Sign Language (ASL) communication for Deaf and Hard of Hearing (DHH) users in Virtual Reality. The study evaluates three mounting positions (head, shoulder, and chest), achieving a high average comprehension success rate of 83.3% with the shoulder-mounted position reaching a peak accuracy of 85%.

TL;DR

Communication for the Deaf and Hard of Hearing (DHH) in Virtual Reality has long been restricted to suboptimal text chats or glitchy avatars. VRSL introduces a pragmatic shift: using body-mounted 360-degree cameras to stream real-life ASL video directly into the VR space. Testing head, shoulder, and chest mounts, the study achieved an 83.3% comprehension rate, proving that video-based signing is not only feasible but preferred by DHH users over traditional text.

The "Uncanny Valley" of VR Accessibility

For DHH users, VR is often a silent, isolating experience. Current SOTA (State-of-the-Art) solutions fall into two categories, both flawed:

  1. Text/Captions: High latency, lacks non-verbal nuance, and is exhausting for "DHH-to-DHH" communication where ASL is the primary language.
  2. Avatars & Hand Tracking: While promising, current hardware has massive "blind spots." If a sign involves touching the face or overlapping hands, the tracking often breaks, leading to "broken" signs that users have to reinvent just to be understood.

The authors' insight was simple yet bold: Stop trying to simulate the signer; just show the signer.

Methodology: Finding the Optimal Vantage Point

The research centered on "Where should the camera go?" A Ricoh Theta V 360° camera was used to capture ASL signs from three distinct first-person perspectives:

  • Head-Mounted: Mimics the signer's own view.
  • Shoulder-Mounted: Offers a stable, slightly offset third-person angle.
  • Chest-Mounted: Provides a central view of both hands but is prone to chin/body obstruction.

Camera Mounting Positions

The study utilized Minimal Pairs—signs that differ by only one parameter (like "Please" vs "Sorry")—to stress-test whether the camera resolution and angle could capture the fine-grained nuances of ASL.

Experimental Battleground: Shoulder vs. The Rest

The results revealed a clear, albeit statistically narrow, winner.

  • Accuracy: The Shoulder Mount led with 85% accuracy. Interestingly, "Location" signs (where the sign happens on the body) were recognized with 100% accuracy, while "Palm Orientation" was the most difficult to discern (80%).
  • Cognitive Load: NASA-TLX data showed the shoulder mount was the least effortful for participants.
  • The Sentence Gap: While individual words were easy to spot (96.67%), complex sentences dropped to 70%. Users reported that the "fish-eye" distortion at the edges of the 360° feed made it hard to track rapid, multi-sign sequences.

Accuracy and NASA-TLX Results

Critical Insight: The Distortion Challenge

The biggest "enemy" identified wasn't the mounting position, but optical distortion. 360-degree cameras use dual fish-eye lenses. Signs performed at the periphery (the edges of the lens) become stretched and blurry. For a language that relies on the precise curl of a finger or the orientation of a palm, this distortion is a significant barrier.

Future Outlook: Beyond the Body-Mount

The VRSL study provides a critical baseline for video-mediated communication in VR. However, the authors and participants agree on a key evolution: moving away from "body-mounted" cameras toward "spatial cameras."

Imagine a small, external 360° camera placed on a desk or a tripod in the user's physical room, streaming a clean, third-person perspective into the VR world. This would eliminate the "bird's-eye" awkwardness of the head mount and the "half-body" visibility issues of the shoulder mount.

Final Takeaway

The DHH community has a strong preference for signing over texting in VR. VRSL proves that video-based communication is the most immediate path to making VR truly inclusive, provided we can solve the "lens distortion" puzzle in future hardware iterations.


Keywords: ASL, Virtual Reality, Human-Computer Interaction, Accessibility, 360-Degree Video.

Find Similar Papers

Try Our Examples

  • Search for recent studies on using multi-camera arrays or de-warping algorithms to reduce peripheral distortion in 360-degree video for gesture recognition.
  • Which paper first proposed the "Chat in the Hat" body-mounted camera system, and how does this VRSL study evolve that concept for immersive virtual environments?
  • Explore research that applies similar video-overlay techniques for sign language communication in Augmented Reality (AR) or telepresence robotics.
Contents
[CHI/ASSETS] VRSL: Breaking the Silence in VR with 360-Degree Sign Language Feeds
1. TL;DR
2. The "Uncanny Valley" of VR Accessibility
3. Methodology: Finding the Optimal Vantage Point
4. Experimental Battleground: Shoulder vs. The Rest
5. Critical Insight: The Distortion Challenge
6. Future Outlook: Beyond the Body-Mount
6.1. Final Takeaway