Imaging the Hidden: Turning Your Smartphone LiDAR into a "Super-Sensor"
Imaging Hidden Objects with Consumer LiDAR via Motion Induced Sampling
This paper introduces a method for Non-Line-of-Sight (NLOS) imaging using consumer-grade LiDAR sensors (<100) found in smartphones. By proposing a Motion-Induced Aperture Sampling (MAS) model and a multi-frame fusion strategy, the authors achieve SOTA-level NLOS 3D reconstruction, multi-object tracking, and camera localization on handheld devices.
TL;DR
Researchers have unlocked the ability to see around corners using the cheap LiDAR sensors found in modern smartphones. By treating handheld motion not as a problem, but as a "synthetic aperture" that gathers more data over time, they've achieved real-time 3D tracking and reconstruction of hidden objects—breaking NLOS imaging out of the lab and into our pockets.
Background: The NLOS Challenge
Non-Line-of-Sight (NLOS) imaging is the "holy grail" of computer vision: seeing objects hidden behind walls by analyzing how light bounces off visible surfaces (the "relay wall"). While research-grade systems use $50,000 sensors, consumer LiDARs are limited by:
- Weak Signals: Eye-safety regulations limit laser power.
- Sparse Data: Low spatial resolution (often only ~100 pixels).
- Dynamics: Handheld motion usually creates blur.

Methodology: Motion-Induced Aperture Sampling (MAS)
The core insight is the MAS Model. Instead of trying to reconstruct a scene from one noisy frame, the authors recognize that a moving camera creates a Synthetic Aperture.
The Math of Motion
Using the Light-Cone Transform (LCT), the authors show that a rigid-body translation of a hidden object corresponds to a simple shift in the measurement space. This allows them to pre-compute a "Canonical STIR" (Space-Time Impulse Response) for an object and then use efficient indexing to compare it against real-time data.
Particle Filtering for Rubustness
Because LiDAR data is noisy, the system uses a Particle Filter. It maintains 1,000 "guesses" (particles) for the hidden object's position. In each frame:
- Propagate: Move particles based on a motion prior (e.g., constant velocity).
- Evaluate: Score particles by "rendering" what the LiDAR should see if the object were at that spot and comparing it to actual data.
- Resample: Keep the best guesses.

Experiments & Results
The authors validated their approach using a smartphone-grade LiDAR across three main tasks:
- 3D Tracking: Tracking moving hidden objects with sub-decimeter accuracy. Even multi-target tracking (e.g., tracking two hands) is possible by clustering the particles.
- 3D Reconstruction: By moving the phone in a scanning motion, they synthesize a large virtual aperture that reveals the 3D shape of a hidden mannequin.
- Camera Localization: In "white-wall" rooms where standard SLAM fails, the system uses hidden objects as landmarks to navigate.

Critical Insights: Why it works
The SOTA performance comes from exploiting priors. By assuming we know the object shape (tracking) or the environment (localization), the "blind deconvolution" problem becomes a much simpler "state estimation" problem. Furthermore, using "multi-bounce light" acts as a natural anti-aliasing filter, allowing sparse sensors to capture more structural information than direct line-of-sight might suggest.
Conclusion & Future Work
This paper represents a paradigm shift toward Plug-and-Play NLOS.
- Democratization: No hours of calibration or expensive gantries.
- Limitation: Currently assumes rigid-body motion and requires knowledge of either shape or environment.
- Outlook: Future iterations could use Machine Learning to learn more complex reflectance models (BRDFs) to work with any natural object, not just retroreflective ones.
The project heralds a future where your vacuum robot can see the cat around the corner, or your AR headset can track your hands even when they are behind your back.
