AMB3R: Bridging the Gap Between Feed-forward Models and Optimization-based SLAM

AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend

2025-01-01
Hengyi Wang, Lourdes Agapito
Summary
Problem
Method
Results
Takeaways
Abstract

AMB3R is a multi-view feed-forward 3D reconstruction model that introduces a sparse, compact volumetric backend to pointmap-based foundation models. It achieves state-of-the-art results in camera pose estimation, metric depth, and dense 3D reconstruction, surpassing traditional optimization-based SLAM and SfM methods.

TL;DR

AMB3R is a breakthrough feed-forward 3D reconstruction model that finally brings "spatial compactness" to the pointmap-based foundation model paradigm. By adding a sparse voxel transformer backend to a frozen front-end, it achieves SOTA performance in Visual Odometry (VO), Structure-from-Motion (SfM), and Metric-scale Reconstruction—all without the need for traditional, time-consuming Bundle Adjustment (BA) or test-time optimization.

Background: The Many-to-One Problem

Modern 3D foundation models like DUSt3R and VGGT have revolutionized the field by regressing Pointmaps—direct mappings from 2D pixels to 3D coordinates. However, these models often treat each pixel independently. In reality, multiple pixels across different views correspond to the same 3D point. Previous models lacked an explicit mechanism to "fuse" these observations into a single, compact 3D representation, leading to geometric noise and drift.

Methodology: The Best of Both Worlds

AMB3R solves this by introducing a Compact Backend on top of a powerful front-end (VGGT).

1. Front-end and Scale Head

The model uses frozen VGGT weights to extract initial pointmaps and geometric features. A new, lightweight Scale Head is trained to regress the metric log-depth of the median pixel, allowing the system to align the reconstruction to real-world units (meters/centimeters).

2. The Sparse Voxel Backend

Instead of letting points float freely, AMB3R aggregates features into a Sparse Voxel Grid.

  • Serialization: These voxels are ordered into a 1D sequence using space-filling curves.
  • Processing: A transformer (Point Transformer v3) performs geometric reasoning within this compact 1D sequence.
  • Injection: The refined features are mapped back to pixels via KNN interpolation and injected into the original decoder using Zero-Convolutions (a technique borrowed from ControlNet). This allows the model to refine geometry while preserving the high-quality confidence functions it learned during pre-training.

Model Architecture

Revolutionary Performance in VO and SfM

One of the most impressive feats of AMB3R is its ability to perform Uncalibrated Visual Odometry. Traditional systems require known camera intrinsics and heavy optimization to prevent drift. AMB3R achieves this in a purely feed-forward manner.

  • Visual Odometry: It maintains a "Keyframe Memory" and performs relative scale alignment without explicit Kabsch–Umeyama alignment, making it robust and fast (~4.2 FPS on an RTX 4090).
  • Structure from Motion: Using a divide-and-conquer image clustering strategy, it scales to large scenes that were previously the exclusive domain of optimization tools like COLMAP.

Experimental Results Comparison Table: AMB3R (Uncalibrated) outperforming traditional DROID-SLAM and other SOTA systems on the TUM benchmark.

Key Results

  • SOTA in Pose & Depth: Leading performance across 13 datasets including NYUv2, KITTI, and ETH3D.
  • SLAM Excellence: Surpassed optimization-based counterparts on TUM and ETH3D Slam Benchmarks for the first time using an uncalibrated feed-forward approach.
  • Efficiency: Only 80 H100 GPU hours are needed for training, making it accessible for academic research compared to the thousands of hours typically required for foundation models.

Qualitative Reconstruction In-the-wild reconstruction of the Longmen Grottoes showcasing the model's generalization capabilities.

Critical Insight: Why This Matters

The core "Aha!" moment of AMB3R is the realization that spatial compactness is not just a storage optimization but a geometric prior. By forcing the network to reason about 3D space in a sparse but unified grid, the model naturally corrects errors that occur in 2D-centric pointmap regression.

Limitations

Despite its success, AMB3R still scales quadratically with the number of images during the attention phase (a common Transformer bottleneck). However, its usage of sparse 3D representation suggests a path toward future models where complexity scales with the "amount of 3D content" rather than the number of pixels.

Conclusion

AMB3R represents a major step toward a unified, generalizable, and scalable feed-forward 3D perception system. By reconciling the efficiency of neural regression with the structural rigor of classical geometry, it sets a new standard for what foundation models can achieve in 3D vision.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use sparse voxel grids or Point Transformers to refine multi-view 3D reconstruction outputs.
  • Which paper first introduced the DUSt3R paradigm of regressing pointmaps, and how does AMB3R structurally modify that decoder architecture?
  • Look for studies applying feed-forward pointmap regression models to large-scale urban reconstruction or autonomous driving datasets like Waymo.
Contents
AMB3R: Bridging the Gap Between Feed-forward Models and Optimization-based SLAM
1. TL;DR
2. Background: The Many-to-One Problem
3. Methodology: The Best of Both Worlds
3.1. 1. Front-end and Scale Head
3.2. 2. The Sparse Voxel Backend
4. Revolutionary Performance in VO and SfM
5. Key Results
6. Critical Insight: Why This Matters
6.1. Limitations
7. Conclusion