Multimodal Models · Vision Foundation Models · AI Agents

Spatial Reasoning in MLLMs: Agent Workspace vs Novel Views vs Pretraining

Spatial reasoning for MLLMs on three axes: SpatialClaw's Python workspace (59.9% on 20 benchmarks), ReRe's synthesized second view (+8.5 VSI-Bench on 2B), and frozen-probe evidence video models carry geometry VLMs lack.

Spatial Reasoning in MLLMs: Agent Workspace vs Novel Views vs Pretraining

Why spatial reasoning is the bottleneck

VLMs got very good at describing images and still fail at questions a child answers from one glance: which object is farther, where the camera moved, whether two pixels in different frames are the same chair. Three papers on this site attack that gap from three different axes, and the useful comparison is between the axes, not just the scores.

SpatialClaw changes the action interface: it gives a VLM a persistent Python kernel with perception and geometry primitives. ReRe changes the inference protocol: it synthesizes a novel viewpoint the camera never captured and lets the model revise its answer. The pretraining-paradigm study changes the representation: it asks whether video-generation models, trained to synthesize scenes, encode more usable 3D structure than VLMs. One fixes how the model acts, one fixes what the model sees, one fixes what the model knows. Your budget and timeline pick the axis.

SpatialClaw: code as the action interface

SpatialClaw’s premise is that spatial agents fail because structured tool calls are too rigid. If a task needs a mask, a depth map, a plane fit, and a trajectory check in sequence, a fixed JSON tool schema forces the interface designer to anticipate every composition. SpatialClaw instead gives the agent a persistent Python workspace: the model writes one cell at a time, calls perception primitives, inspects intermediate outputs, keeps variables alive across steps, and revises when a mask looks wrong.

The results, all training-free on a Gemma 4-31B backbone, are consistent enough to attribute to the harness rather than luck:

Interface (same tools)Avg over 20 spatial benchmarks
No tool access53.4
Single-pass code55.2
Structured tool calls56.7
SpatialClaw persistent workspace59.9

The comparison against the prior spatial-agent baseline in the same table is 59.9 versus 48.7, an 11.2-point gap. On named benchmarks the margins hold: MindCube 72.8 versus 62.4 for structured tool calls and 52.9 for SpaceTools; DSI-Bench 62.9 versus 58.4 and 43.0. The ablation is the part worth copying: removing utility functions barely moves the average (56.9 to 56.4) while removing perception tools drops it to 51.4. The workspace itself, not the size of the tool menu, is doing the work.

ReRe: revise after watching a view that never existed

ReRe (Reason, then Re-reason) is also training-free, but it intervenes on the input instead of the action space. The model first answers a spatial question from an egocentric clip. Then a Geometry-to-Video pipeline reconstructs the scene and renders an elevated, oblique fly-through, and the same model re-watches that synthesized view and confirms or corrects its first answer. No weights change.

On VSI-Bench, the gains are largest for the small models where they matter most:

BackboneBaselineWith ReReDelta
Qwen3-VL-2B22.531.0+8.5
Qwen3-VL-4B30.736.5+5.8
Qwen2.5-VL-7B24.829.5+4.7
InternVL2.5-8B35.536.7+1.2

For scale, GPT-4o sits at 34.0 and Gemini-1.5 Pro at 45.4 on VSI-Bench, so a 2B model with ReRe lands past GPT-4o. The ablations answer the skeptic’s question directly. If the gain came from “thinking twice,” then naively concatenating or interleaving the second view would match it. It does not: single-turn 24.8, concat 25.4, interleaved 25.6, full ReRe 29.5 on Qwen2.5-VL-7B. What earns the points is the structured revise step over a good novel view. Show the model a wrong paired view and the score drops to 23.5, below the no-second-view baseline of 24.8. A bad second view actively degrades the answer. The trajectory ablation agrees: the elevated oblique sweep (29.5) beats a bird’s-eye orbit (27.4) and a mid-level traverse (25.6).

VLM vs video-generation pretraining: who carries the geometry?

The third paper stops asking about test-time machinery and asks what the pretrained backbone already knows. It freezes a batch of vision-language models and video-generation models, trains only a thin probe on top, and reads out three spatial skills: semantic tagging on ScanNet20 (mAP), cross-view instance grouping (T-mIoU), and dense 3D geometry on DL3DV (camera-motion AUC@30, depth AbsRel, point-map error).

TaskVLMs (avg)Video models (avg)Winner
Semantic tagging, ScanNet20 mAP92.0869.89VLM, by 22 points
Instance grouping, T-mIoU22.6613.24VLM
Camera motion, AUC@300.3300.527Video model
Depth, AbsRel (lower better)0.1130.072Video model
Point-map error (lower better)0.2230.152Video model

The split is clean: VLMs own semantics, video-generation models own geometry, because synthesizing video forces a model to internalize where the camera is and how surfaces sit in 3D. The practical answer is the third finding. Concatenate one VGM and one VLM (WAN2.1-T2V-14B plus Qwen3-VL-8B, normalized features), train the same probe, and you get 92.30 mAP and 0.615 camera AUC@30, beating either single backbone on both axes at once. The probe-depth check (rankings stable from probe depth 1 to 6) supports reading the numbers as properties of the frozen representations rather than of the probe.

Key numbers

NumberWhere it comes from
59.9 vs 56.7 vs 53.4SpatialClaw workspace vs structured tool calls vs no tool, 20-benchmark average
72.8 vs 62.4 vs 52.9MindCube: SpatialClaw vs structured tool calls vs SpaceTools
62.9 vs 58.4 vs 43.0DSI-Bench: same three interfaces
22.5 → 31.0 (+8.5)ReRe on VSI-Bench, Qwen3-VL-2B, training-free
24.8 / 25.4 / 25.6 / 29.5 / 23.5ReRe ablation on Qwen2.5-VL-7B: single-turn / concat / interleaved / full ReRe / wrong-view pairing (below baseline)
92.08 vs 69.89 mAPScanNet20 semantic tagging, VLM vs video-model frozen features
0.527 vs 0.330 AUC@30DL3DV camera motion, video models vs VLMs
92.30 mAP, 0.615 AUC@30Naive concat of WAN2.1-T2V-14B + Qwen3-VL-8B features, beats both single backbones

These numbers come from different harnesses and cannot be merged into one leaderboard, and any page that does merge them is lying to you.

When to use which

What the three papers do support is a decision order. Both training-free papers report their largest gains exactly where naive VLMs are weakest, and both cost inference-time compute plus an external tool (a Python kernel, or a geometry-to-video pipeline) rather than a training run. So before budgeting a training job:

  • Failure is multi-step measurement (relative direction, camera motion, counting across views): start with the SpatialClaw-style workspace. It needs a sandboxed Python kernel plus reliable perception primitives, and its own ablation says a small tool set with persistent state beats a large menu of endpoints.
  • Failure is first-glance misjudgment from one ego clip: start with the ReRe-style second viewpoint. It needs a 3D reconstruction and novel-view renderer feeding the same VLM twice, and it degrades gracefully only if the synthesized view is correct, so invest in the geometry pipeline.
  • You are training or fine-tuning anyway: take the geometry from a video-generation backbone. The probe study says you do not need a fusion architecture to start; normalized feature concat already beats both single backbones.

Limits and open questions

Two cautions travel with all three papers. First, every headline here is benchmark spatial reasoning, not embodied competence; SpatialClaw is explicit that it does not prove physically valid planning in the world. Second, gains shrink as backbones get stronger: ReRe adds +8.5 to a 2B model and +1.2 to an 8B one. The interface and viewpoint tricks are biggest wins for the models that could not afford training in the first place.

Open questions the papers leave on the table: whether a persistent workspace and a second viewpoint compose (nobody has run SpatialClaw on top of ReRe-style inputs); whether the VGM-geometry advantage survives full fine-tuning instead of frozen probing; and whether any of these transfers to streaming settings where the model cannot revisit earlier frames.

FAQ

Which method should I try first: SpatialClaw’s code workspace or ReRe’s novel views?

Both papers are training-free, so the cost is engineering, not GPUs. SpatialClaw needs a sandboxed Python kernel plus perception primitives; ReRe needs a 3D reconstruction and novel-view renderer feeding the same VLM twice. If your failure mode is multi-step measurement (relative direction, camera motion, counting across views), start with the workspace. If it is first-glance misjudgment from a single ego clip, start with the second viewpoint.

Does re-watching a second view improve spatial reasoning benchmark scores?

Only with the right protocol. ReRe’s ablation on Qwen2.5-VL-7B: single-turn 24.8, naive concat 25.4, interleaved 25.6, full ReRe 29.5. And pairing the model with a wrong synthesized view scores 23.5, below the 24.8 baseline. The structured revise step over a correct elevated oblique view is what earns the points; a careless second view makes things worse.

Are video-generation models better than VLMs at 3D spatial reasoning?

Split decision, and the split is the point. On the frozen-probe study, video models win geometry (camera AUC@30 0.527 vs 0.330 on DL3DV, depth AbsRel 0.072 vs 0.113) while VLMs win semantics (92.08 vs 69.89 mAP on ScanNet20) and cross-view grouping (22.66 vs 13.24 T-mIoU). A naive concat of both feature sets beats either alone on both axes (92.30 mAP, 0.615 AUC@30).

Do these methods prove MLLMs understand 3D?

No, and SpatialClaw says so itself: the improvement should be attributed to the harness plus tools, not to the base model. All three papers improve how spatial evidence reaches the model or how it is computed over. None claims a native 3D world model inside the VLM.

Verdict in one line: fix the interface or the viewpoint before retraining, and if you pretrain, take the geometry from video-generation features. The details with numbers are in the SpatialClaw, ReRe, and pretraining-paradigm paper pages.