Imaging Hidden Objects with Consumer LiDAR: The Dawn of Plug-and-Play NLOS

Imaging Hidden Objects with Consumer LiDAR via Motion Induced Sampling

2026-01-01
Siddharth Somasundaram, Aaron Young, Akshat Dave, Adithya Pediredla, Ramesh Raskar
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a framework for Non-Line-of-Sight (NLOS) imaging using smartphone-grade, low-cost LiDAR sensors. By leveraging a novel Motion-Induced Aperture Sampling (MAS) model and multi-frame fusion, the authors achieve 3D reconstruction, multi-object tracking, and camera localization on hardware costing less than $100.

TL;DR

Researchers from MIT and Dartmouth have broken the "lab-only" barrier of Non-Line-of-Sight (NLOS) imaging. By using a smartphone-grade LiDAR and a clever multi-frame fusion strategy called Motion-Induced Aperture Sampling (MAS), they can see around corners in real-time. This turns everyday devices into high-tech periscopes capable of 3D tracking and self-localization using hidden landmarks.

The Problem: Why Your Phone Can't See Around Corners (Yet)

NLOS imaging is essentially trying to use a wall as a mirror. However, a painted wall is a diffuse surface—it scatters light in every direction. To reconstruct a hidden object from these faint, scattered reflections, you typically need:

  1. High Power: To catch the few photons that make the triple-bounce trip (Laser → Wall → Object → Wall → Sensor).
  2. High Resolution: To sample the "virtual aperture" on the wall densely.
  3. Stability: To prevent motion blur.

Consumer LiDARs, like those in iPhones or Roombas, fail on all three counts. They have low power (for eye safety), low resolution (~100 pixels), and are almost always moving in a user's hand.

The Insight: Motion is a Feature, Not a Bug

The authors propose that instead of fighting camera and object motion, we should use it. Inspired by Burst Photography (which stacks noisy images to get a clean one) and Synthetic Aperture Radar (SAR) (which uses movement to simulate a larger antenna), they developed the MAS Model.

Methodology: The MAS Model

The core mathematical engine is the Light-Cone Transform (LCT). The MAS model demonstrates that in the LCT-transformed space, the relationship between an object and its measurement is a 3D convolution.

Crucially, they prove that a rigid-body translation of an object results in a simple shift of the Space-Time Impulse Response (STIR). This allows the system to decouple:

  • Object Shape: A time-independent "canonical" STIR.
  • Motion: A time-dependent shift/translation.
  • Camera Pose: A spatial sampling function.

MAS Model Architecture

Real-Time Inference via Particle Filtering

Solving for shape, motion, and pose simultaneously is "highly non-convex" (translation: computationally nightmarish). The authors circumvent this by observing that in most practical tasks, two out of three variables are usually known:

  • Tracking: Known shape/pose, solve for object position.
  • Localization: Known shape/object position, solve for camera pose.

They use a Particle Filter to track these variables. By maintaining 1,000 "guesses" (particles) and weighting them against real-time LiDAR data using a dot-product score function, the system can "lock on" to a hidden object even when individual frames are mostly noise.

Particle Filtering Process

Experimental Results: Seeing is Believing

The team tested their approach using an off-the-shelf ST VL53L8CX sensor (retailing for under $100).

  1. Object Tracking: They achieved single and multi-object tracking with a mean error of ~4.7 cm.
  2. Diffuse Objects: While retroreflective materials (like safety vests) yield the best results, they proved the system works on normal diffuse objects (like a mannequin), albeit with lower SNR.
  3. Camera Localization: In a room with white, featureless walls where standard Visual Odometry (like ARKit) would fail, their system used a hidden object around the corner as a stable landmark to track the camera's path.

Quantitative Results

Critical Analysis & Future Outlook

The Takeaway: This research democratizes NLOS. We are moving away from $100k lab setups toward "plug-and-play" software solutions.

Limitations:

  • Reflectance: The model still struggles with complex, non-homogeneous BRDFs (materials that reflect light in weird ways).
  • Priors: It requires a "canonical shape" (knowing what the hidden thing looks like) to track effectively in real-time.

Future Work: The authors suggest using Machine Learning to replace handcrafted score functions, potentially allowing the system to learn how to see even more complex hidden environments without predefined shapes.

Ultimately, this work suggests that your future robot vacuum might not just avoid the chair it sees—it will avoid the cat running toward it from the other room.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize SPAD (Single-Photon Avalanche Diode) sensors in handheld mobile devices for 3D scene understanding or occluded object detection.
  • Which paper first proposed the Light-Cone Transform (LCT) for NLOS imaging, and how does the current MAS model extend the LCT to handle dynamic joint motion of both the camera and the object?
  • Are there any studies that apply Particle Filtering or Monte Carlo Localization to time-of-flight (ToF) data for indoor robot navigation where line-of-sight visibility is limited?
Contents
Imaging Hidden Objects with Consumer LiDAR: The Dawn of Plug-and-Play NLOS
1. TL;DR
2. The Problem: Why Your Phone Can't See Around Corners (Yet)
3. The Insight: Motion is a Feature, Not a Bug
3.1. Methodology: The MAS Model
4. Real-Time Inference via Particle Filtering
5. Experimental Results: Seeing is Believing
6. Critical Analysis & Future Outlook