VLM3: Vision Language Models Are Native 3D Learners

VLM3: Vision Language Models Are Native 3D Learners

2026-05-01
Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu, Vikas Chandra, Yangyang Shi
Summary
Problem
Method
Results
Takeaways
Abstract

VLM3 is a scalable framework that demonstrates standard Vision Language Models (VLMs) can become native 3D learners without task-specific architectures or loss functions. By leveraging focal length unification, text-based pixel normalization, and optimized data scaling, VLM3-4B achieves SOTA or competitive results in depth estimation, pixel correspondence, and camera pose estimation.

TL;DR

VLM3 proves that you don't need complex geometric decoders or regression losses to master 3D vision. By simply "talking" to the model about pixels in a normalized coordinate space and unifying camera intrinsics through resizing, a standard 4B VLM can outperform billion-parameter expert models in depth, pose, and correspondence estimation.

Background: Beyond Semantic Labels

While Vision Language Models (VLMs) have conquered semantic understanding (e.g., "What is in this image?"), 3D understanding has remained the fortress of "Expert Models." These experts rely on task-specific headers and complex physics-based losses. VLM3 strikes at the heart of this complexity, arguing that 3D is just another modality that can be learned through text-based scaling.

The Three Pillars of VLM3

The authors identify that standard VLMs fail at 3D not because of their architecture, but because of how the data is presented. They introduce three non-invasive fixes:

  1. Resolving Camera Ambiguity: Since a small object close-up looks the same as a large object far away, VLM3 resizes images to unify the focal length to 1000 pixels. This allows the model to learn a consistent mapping between pixel size and metric depth.
  2. Normalized Textual Grounding: Instead of drawing markers on images (visual prompting), VLM3 uses text like (x, y) coordinates normalized to [0, 2000). This makes the model more efficient, as it can answer multiple questions about one image in a single forward pass.
  3. The Power of Mixture: The secret sauce isn't a new layer; it's the weight of the datasets. Scaling the training to 32M samples with 320M labeled pixels proved that data volume and balanced mixtures are the primary drivers of 3D competence.

VLM3 Architecture and Task Overview

Crushing the "Regression" Myth

Perhaps the most shocking finding is that regression losses are unnecessary. Traditionally, predicting a camera's rotation or a pixel's depth required Mean Squared Error (MSE) or similar continuous losses. VLM3 treats these as text tokens (Classification).

  • Translation: "The camera moves right, unit vector (0.5, 0.1, 0.8)."
  • Depth: "The pixel distance is 5.42 meters."

By converting geometry into a vocabulary, VLM3 achieves a 94% accuracy in pose estimation, matching "Giant" models that use heavy geometric supervision.

Experimental Performance

The VLM3-4B model was tested against both generalist VLMs and 3D experts:

  • Depth Estimation: Surpassed DepthLM-7B, reaching a δ1 of 0.90.
  • Correspondence: Outperformed DKM and RoMa, reducing error by 10x over Qwen3-VL baselines.
  • Object Understanding: Beat SpatialRGPT-8B without needing any specialized region encoders.
TaskVLM3-4B (Ours)Expert SOTA
Depth (Average δ1)0.9040.945 (UniDepthV2)
Pose (AUC@30°)94.094.7 (DA3-Giant)
Correspondence (EPE)15.377.89 (UFM)

Experimental Results Comparison

Deep Insight: Is Scaling Always Better?

An interesting "Ablation Study" in the paper reveals that Model Scaling is not yet the bottleneck. Training an 8B or 32B model actually resulted in lower accuracy compared to the 4B model on the same 26M image dataset. This suggests that for 3D tasks, the quality and variety of 3D-grounded data are currently more important than the parameter count. We are still in the "data-limited" regime for 3D VLM training.

Conclusion & Future Impact

VLM3 marks a "bitter lesson" moment for 3D vision. It suggests that specialized geometric modules might eventually be replaced by general-purpose tokens. For developers and researchers, the takeaway is clear: focus on camera calibration and data normalization, and let the VLM's inherent pattern-matching capabilities do the rest.

Limitations: While VLM3 is efficient, it still lags slightly behind the absolute best-in-class specialized models in dense correspondence. However, the gap is closing rapidly through data scaling alone.

Find Similar Papers

Try Our Examples

  • Which recent papers explore the transition from geometric regression losses to token-based classification for 3D vision tasks in Large Multimodal Models?
  • What is the origin of focal length unification as a strategy for solving monocular depth ambiguity, and how does VLM3's implementation compare to the original DepthLM?
  • Are there studies investigating how the 2D spatial biases of pre-trained VLM encoders (like Qwen-VL or CLIP) facilitate or hinder the learning of 3D geometric properties?
Contents
VLM3: Vision Language Models Are Native 3D Learners
1. TL;DR
2. Background: Beyond Semantic Labels
3. The Three Pillars of VLM3
4. Crushing the "Regression" Myth
5. Experimental Performance
6. Deep Insight: Is Scaling Always Better?
7. Conclusion & Future Impact