SPACENUM: Why Your VLM Can't Actually Do Spatial Math
SPACENUM: Revisiting Spatial Numerical Understanding in VLMs
This paper introduces SPACENUM, a unified framework designed to evaluate Vision-Language Models (VLMs) on their ability to ground numerical values in spatial perception across two settings: dynamic transitions (exploration) and static layouts (reasoning). It establishes bidirectional tasks, NUM2SPACE and SPACE2NUM, revealing that current SOTA models like Qwen2.5-VL and InternVL3.5 still struggle significantly with precise spatial numerical understanding.
TL;DR
Despite the impressive progress in Vision-Language Models (VLMs), they remain "spatial-numerical illiterates." A new benchmark, SPACENUM, reveals that even top-tier models like Qwen2.5-VL and InternVL3.5 perform only slightly better than random guessing when tasks require mapping precise numbers to spatial changes. The core issue isn't a lack of reasoning—it's a lack of structured spatial abstraction.
The "Illusion" of Spatial Understanding
When a robot agent says "Rotate Left 20°," we assume it understands the magnitude of that 20°. However, current research suggests this is often a "shallow cue" fallback. Previous benchmarks measured topology (is A next to B?), but SPACENUM digs into the metric properties: Is the model actually grounded in the numbers?
The authors identify two ways numbers live in space:
- Dynamic Transitions: Numbers as the magnitude of movement (e.g., moving forward 0.4m).
- Static Layouts: Numbers as coordinates (x, y, z) that define the "where" and "how big" of objects in a scene.
Methodology: The Bidirectional Stress Test
The researchers developed two complementary tasks across these settings:
- NUM2SPACE: Given an image and a number (e.g., "move 0.6m"), identify the correct resulting image.
- SPACE2NUM: Given two images (before and after), derive the numerical magnitude of the change.
Figure 1: The dual-direction framework of SPACENUM testing mapping from vision to numbers and vice-versa.
Key Insights: Why VLMs Fail
The results were sobering. The best-performing model, Qwen2.5-VL-72B, achieved only 39.8% accuracy, where a random guess would yield 30%.
1. The "Reasoning" Paradox
Common wisdom suggests that "Chain-of-Thought" (CoT) or "Thinking" models solve complex problems. In SPACENUM, enabling reasoning provided marginal to zero gain. Traces revealed that models often stop at coarse cues (e.g., "The chair moved left") but fail the fine-grained comparison required to distinguish 10° from 30°.
2. Modality Asymmetry
Models are consistently better at SPACE2NUM (interpreting visual shifts into numbers) than NUM2SPACE (imagining a visual outcome from a number). This suggests VLMs behave more like "describers" than "simulators."
3. The Geometry Problem
The study found a lack of Rotational Symmetry. If you ask a model to rotate left 20°, it gives one answer; if you ask it to rotate right 340° (the same physical outcome), the performance collapses. The internal representations are not geometrically invariant.
Table 1: Comprehensive performance across 18 models. Note how close most scores are to the Random Guess baseline.
Can We Fix It?
The authors tried several interventions:
- Visual Anchors: Adding cues to help measure distance. (Result: Minimal impact).
- Numerical Reformulation: Changing "0.2m" to "20cm". (Result: Minimal impact).
- Structured Abstraction: Replacing raw images with 2D/3D boxes. (Result: Major Improvement!)
This last point is critical. It proves that the bottleneck isn't the model's "brain" (the LLM part), but its "eyes" (the Vision Encoder's inability to extract structured coordinate frames from raw pixels).
Future Outlook: Moving Beyond Pixels
As we push toward "Physical AI" and embodied agents, the "shallow spatial sensitivity" documented in SPACENUM must be addressed. Fine-tuning on 1D spatial data was shown to partially transfer to 3D, suggesting that spatial numerical understanding is a learnable skill, but one that requires a specific "data recipe" (roughly 75% layout data and 25% transition data).
Closing Thought: If we want VLMs to actually command robots, we need to stop teaching them to just "see" and start teaching them to "measure."
Figure 2: Analysis showing that as models scale, they make "better" mistakes (closer to the truth), but precise grounding remains elusive.
