Smartphone LiDAR’s "X-Ray Vision": Seeing Around Corners via Motion-Induced Sampling
Imaging Hidden Objects with Consumer LiDAR via Motion Induced Sampling
The paper introduces a method for non-line-of-sight (NLOS) imaging using consumer-grade smartphone LiDAR. By leveraging a novel Motion-induced Aperture Sampling (MAS) model and multi-frame fusion, the authors achieve 3D reconstruction, multi-object tracking, and camera localization on low-cost hardware (< $100) that previously required expensive research-grade sensors.
TL;DR
Researchers from MIT and Dartmouth have unlocked the ability to see around corners using the standard LiDAR sensors found in modern smartphones. By treating the natural shaking of a handheld camera and the movement of objects as a "synthetic aperture," their MAS (Motion-induced Aperture Sampling) model overcomes the low-quality signals of consumer hardware to track hidden objects and map hidden rooms in real-time.
Research Positioning: This is a pivotal "democratization" paper. It moves Non-Line-of-Sight (NLOS) imaging out of the 100 consumer chip, replacing expensive hardware precision with clever algorithmic fusion.
The Barrier: Why Your Phone Can't (Usually) See Around Corners
NLOS imaging works by using a wall as a "virtual mirror." We fire a laser at the wall, it bounces to a hidden object, returns to the wall, and finally hits our sensor.
While research-grade setups use High-Power Lasers and picosecond-precision SPADs, consumer LiDAR (like on an iPhone or AR headset) faces three "deal-breakers":
- Low SNR: To keep lasers "eye-safe," power is kept extremely low.
- Poor Resolution: Consumer sensors have very few pixels (often just ~100).
- Dynamic Chaos: In the real world, both the camera and the object are moving, creating a blurred mess of temporal data.
Methodology: The MAS Model and Particle Filtering
The core innovation is the Motion-induced Aperture Sampling (MAS) Model. The authors observed that if you transform the raw time-of-flight data using the Light-Cone Transform (LCT), the complex physics of light bouncing around a corner becomes a simple 3D convolution.
1. Decoupling the "What" from the "Where"
The MAS model breaks down the measurements into three components:
- Object Shape: A canonical, time-independent representation of the hidden item.
- Object Motion: A simple translation of that representation in space-time.
- Camera Pose: A spatial sampling mask determined by where you point your phone.

2. Particle Filtering: Turning Noise into Probability
Instead of trying to "solve" a single blurry frame, the authors use a Particle Filter. They maintain 1,000 "guesses" (particles) for the hidden object's position. As the camera moves, it gathers new viewpoints. Particles that don't match the new data are killed off; those that match are duplicated. This allows the system to build SNR over time, much like "Burst Photography" on your phone's night mode.
Experiments: Tracking and Mapping
The team tested their method on a smartphone-grade sensor with only 100 pixels.
- 3D Tracking: They tracked moving hidden objects with a precision of ~4.7 cm. Even when the object was diffuse (not shiny), the system maintained a probabilistic "map" of its location.
- Camera Localization: Imagine a robot in a white, featureless hallway. It can't use standard vision to see where it is. By looking at a hidden object around the corner, this system uses that object as a "landmark" to calculate the robot's own position.
- Multi-Object Tracking: Using K-means clustering on the particles, the system could distinguish between multiple hidden people or objects simultaneously.

Critical Insight: Opportunistic Blur as Anti-Aliasing
One of the most profound takeaways in the paper is the idea of Multi-Bounce Light as Opportunistic Blur. In standard imaging, low-pixel sensors suffer from "aliasing" (jagged edges). However, because NLOS light is naturally blurred by the wall (the PSF kernel ), the wall acts as a natural anti-aliasing filter. This means we can actually recover finer details of a hidden object's rough shape than we could if we were looking at it directly with the same low-res sensor.

Conclusion & Future Work
This research marks the transition of NLOS from an optical curiosity to a functional feature for consumer electronics.
Limitations: Currently, the model assumes basic rigid-body translation (no rotations). It also works best with retroreflective materials (like high-vis vests), though it demonstrated success with diffuse objects.
The Road Ahead: The authors have demonstrated that a $5 ST VL53L8CX sensor (found in many consumer devices) is sufficient for NLOS tracking. We are likely years, not decades, away from autonomous vacuum cleaners and AR glasses that "know" a person is walking toward a doorway before they even come into view.
