UbiMouse: Transforming Public Displays into Touchless Interfaces for a Post-Pandemic World

Sustainable society with a touchless solution using UbiMouse under the pandemic of COVID-19

2021-08-05
Daisuke Akagawa, Junichi Takatsu, Ryoji Otsu, Seiichi Hayashi, Benjamin Vallet
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces UbiMouse, an AI-driven touchless interface solution designed to convert standard contact-based devices (ATMs, restaurant kiosks) into contactless ones using computer vision. Utilizing a ResNet34-based convolutional backbone followed by a regression model, it enables high-accuracy cursor control via finger movements in the air, captured by standard webcams.

TL;DR

UbiMouse is an AI software solution that eliminates the need for physical contact with public screens like ATMs and kiosks. By combining a ResNet34 feature extractor with a regression network, it translates simple finger movements captured by a standard webcam into precise mouse cursor coordinates. It offers a hygienic, low-cost alternative to traditional touch panels, even functioning when users are wearing gloves.

Background and Motivation: The Hygiene Imperative

Before 2020, touchscreens were the gold standard for intuitive public interaction. However, the COVID-19 pandemic reframed these surfaces as potential vectors for pathogen transmission. Beyond hygiene, "touch sensing faults" occur frequently in industrial or medical settings where users must wear thick gloves.

The authors identified a gap: while specialized sensors like Leap Motion exist, there is a massive need for a software-first approach that uses ubiquitous hardware (standard webcams) to retrofit existing infrastructure.

Methodology: From Vision to Coordinates

The core of UbiMouse is its ability to map a 3D physical gesture (pointing in the air) onto a 2D digital coordinate system .

1. The Model Architecture

UbiMouse leverages a dual-stage neural network architecture:

  • Feature Extraction: A ResNet34 backbone (pretrained on ImageNet) processes the 480x640 input frames. The authors removed the final 1000-node classification layer to obtain a high-dimensional feature tensor (512 channels).
  • Position Regression: To convert these features into coordinates, the system applies both Average Pooling and Max Pooling. These are concatenated into a 1024-dimensional vector, which passes through a fully connected regression network with Dropout (20%) and Batch Normalization.

Overall Model Architecture Figure 1: The UbiMouse pipeline, showing the transition from image input to coordinate output.

2. Implementation Strategies

The training utilized a human-in-the-loop dataset creation process. Users pointed at specific dots on a screen while a webcam captured their finger positions, resulting in 1,416 high-quality labeled data points. The authors applied Cyclical Learning Rates to find an optimal learning rate (settling on ), ensuring stable convergence and preventing the model from getting stuck in local minima.

Regression Part Structure Figure 2: Detailed view of the regression head designed for coordinate estimation.

Experimental Results

The system was evaluated using Mean Squared Error (MSE) between predicted and true pixel coordinates. The training loss curve indicates a smooth convergence, suggesting the ResNet34 features were highly representative of the pointing gesture.

  • User Feedback: Real-world tests in restaurants and signage environments showed that the mouse pointer followed finger movements with minimal latency.
  • Hardware Versatility: The system was successfully deployed both as a standalone USB sensor device and as background software for existing PCs.

Loss Convergence Figure 3: Training loss transition demonstrating the effectiveness of the chosen learning parameters.

Critical Insight: Why This Works

The brilliance of UbiMouse lies in its simplicity. Instead of performing full skeletal hand tracking (which is computationally expensive and prone to jitter), it treats the problem as a direct spatial regression task. By focusing only on the relationship between the image appearance and the target coordinate, the model bypasses the need for complex inverse kinematics.

Future Outlook and Limitations

While UbiMouse is highly effective, the authors note that quantitative evaluation on "untraceable" public users remains a challenge. Future work will focus on:

  1. Model Compression: Reducing the weight of the model for deployment on low-power edge devices.
  2. Robustness: Creating more diverse test cases to handle varying lighting conditions and background clutter.

UbiMouse marks a significant step toward "Sustainable Societies" by providing a technical solution that balances the convenience of digital interfaces with the safety requirements of a global health crisis.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize ResNet-based regression for precise 2D coordinate mapping in human-computer interaction tasks.
  • Which original research first introduced the Cyclical Learning Rate (CLR) method, and how does the LR range test used in this paper improve model convergence?
  • Investigate studies that have extended touchless finger-tracking systems to multi-modal environments involving voice or gesture-based selection confirmation.
Contents
UbiMouse: Transforming Public Displays into Touchless Interfaces for a Post-Pandemic World
1. TL;DR
2. Background and Motivation: The Hygiene Imperative
3. Methodology: From Vision to Coordinates
3.1. 1. The Model Architecture
3.2. 2. Implementation Strategies
4. Experimental Results
5. Critical Insight: Why This Works
6. Future Outlook and Limitations