SPACENUM: Why Your VLM Can't Actually Do Spatial Math

SPACENUM: Revisiting Spatial Numerical Understanding in VLMs

2026-01-01
Jianshu Zhang, Yijiang Li, Huifeixin Chen, Haoran Lu, Letian Xue, Bingyang Wang, Han Liu
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces SPACENUM, a unified framework designed to evaluate Vision-Language Models (VLMs) on their ability to ground numerical values in spatial perception across two settings: dynamic transitions (exploration) and static layouts (reasoning). It establishes bidirectional tasks, NUM2SPACE and SPACE2NUM, revealing that current SOTA models like Qwen2.5-VL and InternVL3.5 still struggle significantly with precise spatial numerical understanding.

TL;DR

Despite the impressive progress in Vision-Language Models (VLMs), they remain "spatial-numerical illiterates." A new benchmark, SPACENUM, reveals that even top-tier models like Qwen2.5-VL and InternVL3.5 perform only slightly better than random guessing when tasks require mapping precise numbers to spatial changes. The core issue isn't a lack of reasoning—it's a lack of structured spatial abstraction.

The "Illusion" of Spatial Understanding

When a robot agent says "Rotate Left 20°," we assume it understands the magnitude of that 20°. However, current research suggests this is often a "shallow cue" fallback. Previous benchmarks measured topology (is A next to B?), but SPACENUM digs into the metric properties: Is the model actually grounded in the numbers?

The authors identify two ways numbers live in space:

  1. Dynamic Transitions: Numbers as the magnitude of movement (e.g., moving forward 0.4m).
  2. Static Layouts: Numbers as coordinates (x, y, z) that define the "where" and "how big" of objects in a scene.

Methodology: The Bidirectional Stress Test

The researchers developed two complementary tasks across these settings:

  • NUM2SPACE: Given an image and a number (e.g., "move 0.6m"), identify the correct resulting image.
  • SPACE2NUM: Given two images (before and after), derive the numerical magnitude of the change.

Experimental Framework Figure 1: The dual-direction framework of SPACENUM testing mapping from vision to numbers and vice-versa.

Key Insights: Why VLMs Fail

The results were sobering. The best-performing model, Qwen2.5-VL-72B, achieved only 39.8% accuracy, where a random guess would yield 30%.

1. The "Reasoning" Paradox

Common wisdom suggests that "Chain-of-Thought" (CoT) or "Thinking" models solve complex problems. In SPACENUM, enabling reasoning provided marginal to zero gain. Traces revealed that models often stop at coarse cues (e.g., "The chair moved left") but fail the fine-grained comparison required to distinguish 10° from 30°.

2. Modality Asymmetry

Models are consistently better at SPACE2NUM (interpreting visual shifts into numbers) than NUM2SPACE (imagining a visual outcome from a number). This suggests VLMs behave more like "describers" than "simulators."

3. The Geometry Problem

The study found a lack of Rotational Symmetry. If you ask a model to rotate left 20°, it gives one answer; if you ask it to rotate right 340° (the same physical outcome), the performance collapses. The internal representations are not geometrically invariant.

Performance Gap Table 1: Comprehensive performance across 18 models. Note how close most scores are to the Random Guess baseline.

Can We Fix It?

The authors tried several interventions:

  • Visual Anchors: Adding cues to help measure distance. (Result: Minimal impact).
  • Numerical Reformulation: Changing "0.2m" to "20cm". (Result: Minimal impact).
  • Structured Abstraction: Replacing raw images with 2D/3D boxes. (Result: Major Improvement!)

This last point is critical. It proves that the bottleneck isn't the model's "brain" (the LLM part), but its "eyes" (the Vision Encoder's inability to extract structured coordinate frames from raw pixels).

Future Outlook: Moving Beyond Pixels

As we push toward "Physical AI" and embodied agents, the "shallow spatial sensitivity" documented in SPACENUM must be addressed. Fine-tuning on 1D spatial data was shown to partially transfer to 3D, suggesting that spatial numerical understanding is a learnable skill, but one that requires a specific "data recipe" (roughly 75% layout data and 25% transition data).

Closing Thought: If we want VLMs to actually command robots, we need to stop teaching them to just "see" and start teaching them to "measure."

Error Analysis Figure 2: Analysis showing that as models scale, they make "better" mistakes (closer to the truth), but precise grounding remains elusive.

Find Similar Papers

Try Our Examples

  • Search for recent papers that focus on metric-grounded spatial reasoning or quantitative coordinate estimation in Vision-Language Models.
  • Which studies first introduced the concept of "Cognitive Maps" in the context of Large Multimodal Models, and how does SPACENUM's approach to static layouts build upon them?
  • Find research that investigates the effectiveness of Reinforcement Learning with Graded Rewards for improving spatial grounding in embodied AI agents.
Contents
SPACENUM: Why Your VLM Can't Actually Do Spatial Math
1. TL;DR
2. The "Illusion" of Spatial Understanding
3. Methodology: The Bidirectional Stress Test
4. Key Insights: Why VLMs Fail
4.1. 1. The "Reasoning" Paradox
4.2. 2. Modality Asymmetry
4.3. 3. The Geometry Problem
5. Can We Fix It?
6. Future Outlook: Moving Beyond Pixels