A benchmark for spatial intelligence
Same Space or Not
Benchmarking Spatial Intelligence
Through Cross-View Contradictions
Four views. One contradiction.
Three images agree on one unchanged 3D scene. A fourth does not.
Can a vision-language model find it and explain why?
See the challenge
Which view breaks the shared scene?
Explore one representative question for each edit type. Look across all four views, then select the inconsistent image.
These examples come from the paper. Edit labels are shown here for exploration. In the standard evaluation, models receive neither the edit label nor a marked target. Your selections stay in this page and are not submitted.
The benchmark
A contradiction must be attributable.
Finding two disagreeing images is not enough. With no trusted reference view and no marked target, the model must identify which view violates the scene supported by the others.
Scroll sideways to view all four charts, or tap the figure to open it at full size.
iTHOR and ProcTHOR are reported jointly for source counts because their room configurations and components overlap in the benchmark. RoboTHOR is reported separately.
Validity-guided constraint discovery
Review of more than 10,000 development candidates turned failure cases into generation constraints. Each accepted item preserves diagnostic evidence in the intruder and at least two anchors while excluding tested 2D shortcuts.
- 01
Check the scene and evidence
Verify physical integrity, the recorded edit and clear visual evidence using the images, answer key and top-down views.
- 02
Verify a unique answer
A fresh reviewer first solves the four-image question without the key, records evidence and ambiguity, then checks for alternative answers.
- 03
Reject visual shortcuts
Exclude questions solvable through isolated appearance, object inventories or nearly aligned pixel comparisons.
Full benchmark results
A wide gap remains.
Six VLMs are evaluated on the same 1,000 questions. Switch between selecting the right view and supporting that selection with valid evidence.
Answer accuracy is the percentage of questions with the correct final choice.
| GPT-6 Astra Pro | 39.70 | 92.08 | 27.20 | 7.23 | 64.36 | 61.39 | 83.17 |
|---|---|---|---|---|---|---|---|
| GPT-5.6 Sol Pro | 24.20 | 58.42 | 14.40 | 6.36 | 24.75 | 42.57 | 56.44 |
| Gemini 3.5 Flash | 20.60 | 46.53 | 13.20 | 9.54 | 19.80 | 30.69 | 41.58 |
| Claude Fable 5 | 19.30 | 39.60 | 14.00 | 5.20 | 24.75 | 31.68 | 42.57 |
| Gemini 3.6 Flash | 18.70 | 34.65 | 16.00 | 5.78 | 23.76 | 29.70 | 37.62 |
| Grok 4.5 | 14.30 | 22.77 | 12.00 | 7.80 | 23.76 | 22.77 | 15.84 |
| Human | 83.30 | 97.03 | 71.60 | 76.01 | 95.05 | 98.02 | 97.03 |
| Random choice | 25.00 | 25.00 | 25.00 | 25.00 | 25.00 | 25.00 | 25.00 |
Bold values mark the best VLM in each category. Human participants supplied one answer per question, yielding 833 correct choices. Human explanation scores were not measured. Seventeen non-generative vision encoders achieved 16.9% to 21.7% answer accuracy.
Rotation is particularly difficult. Even the best VLM answer accuracy in this category is 9.54%, compared with 76.01% for humans.
What SSoN reveals
Key findings
Beyond the performance gap, three findings show where models struggle across views and what helps them improve.
Six VLMs · 1,000 questions eachCorrect choices can hide unsupported explanations.
404 of 1,368 correct selections lack sufficient valid evidence. True solve reveals what answer accuracy alone misses.
29.5%of correct selections
lack valid support
Explore the evidence
Correct choices can hide unsupported explanations.
404 of 1,368 correct selections lack sufficient valid evidence. True solve reveals what answer accuracy alone misses.
lack valid support
Accuracy decomposed
Right answer. Right reason?
Questions (%)
| Model | Answer accuracy (%) | True solve (%) | Partly guessed (%) |
|---|---|---|---|
| GPT-6 Astra Pro | 39.7 | 34.9 | 4.8 |
| GPT-5.6 Sol Pro | 24.2 | 16.9 | 7.3 |
| Gemini 3.5 Flash | 20.6 | 11.6 | 9.0 |
| Claude Fable 5 | 19.3 | 13.1 | 6.2 |
| Gemini 3.6 Flash | 18.7 | 11.9 | 6.8 |
| Grok 4.5 | 14.3 | 8.0 | 6.3 |
A factual explanation must give sufficient evidence for the chosen view. A correct letter alone is not enough.
Partly guessed includes fabricated changes and insufficient evidence. These cases are not all observed hallucinations.
Every correct response was reviewed by humans.
Ten Claude Opus 5 reviewer seats assisted with explanation screening. Every correct-choice response received an independent, agent-blind human judgment followed by full evidence review. Final labels are human decisions.
100 fixed diagnostic questionsGeometric support strengthens contradiction resolution.
The reconstruction workflow improves accuracy by 23 points and True solve by 22. Supplied camera geometry achieves the highest scores.
61%accuracy with
camera geometry
Explore the evidence
Geometric support strengthens contradiction resolution.
The reconstruction workflow improves accuracy by 23 points and True solve by 22. Supplied camera geometry achieves the highest scores.
camera geometry
Prompting, reconstruction and geometry
What helps models find the contradiction?
Questions (%)
| Condition | Answer accuracy (%) | True solve (%) |
|---|---|---|
| Base | 24 | 12 |
| Comparison | 40 | 23 |
| Scene description | 37 | 29 |
| Edit taxonomy | 44 | 32 |
| Comparison + taxonomy | 48 | 34 |
| Reconstruction | 47 | 34 |
| Camera geometry | 61 | 48 |
over Base
over Base
with camera geometry
Adding calibration to comparison prompting raises both move scores from 20% to 70%. All 14 correct move selections have valid support.
Even with camera geometry, rotation reaches 38.89% accuracy and 22.22% True solve. Reconstruction yields only two move and one rotate True solve responses.
What each condition adds
Comparison aligns landmarks, instances, expected visibility and corroborating views.
Scene description adds a written account of the 3D arrangement before deciding.
Edit taxonomy supplies all six possible edit categories without revealing the current edit.
Comparison + taxonomy combines the comparison strategy and category knowledge.
Reconstruction executes and inspects a sparse spatial 3D reconstruction without supplied geometry.
Camera geometry adds camera intrinsics, extrinsics and image dimensions to comparison prompting.
A deliberately selected diagnostic set
The set contains exactly 12 original True solve responses, 12 original Partly guessed responses and 76 wrong responses from Astra Pro. True solve and Partly guessed counts match within each edit category. Every category contributes at least 12 questions and all seven sources are represented.
Base reuses the original responses. Other conditions are new runs. These comparisons describe the selected diagnostic set rather than an average improvement over the full benchmark.
Recon uses GPT-6 Astra at the highest of six available reasoning-effort settings. Humans verify the executed spatial reconstruction, reconstructed camera views and its use in reasoning. Runs lacking execution or required evidence were repeated. The first compliant response was retained regardless of correctness. The reported gains compare the complete reconstruction workflow with Base.
Qwen2.5-VL-7B · 201 held-out questions · Unseen roomsTargeted supervision improves performance on unseen rooms.
Accuracy rises from 21.89% to 33.33% with answer supervision and 41.79% with rationale and answer supervision.
+19.90accuracy points
over zero-shot
Explore the evidence
Targeted supervision improves performance on unseen rooms.
Accuracy rises from 21.89% to 33.33% with answer supervision and 41.79% with rationale and answer supervision.
over zero-shot
Qwen2.5-VL-7B-Instruct
Learning to identify the contradiction
Answer accuracy (%)
| Supervision | Accuracy (%) | Correct | 95% Wilson confidence interval |
|---|---|---|---|
| Zero-shot | 21.89 | 44 / 201 | [16.7, 28.1] |
| Answer supervision | 33.33 | 67 / 201 | [27.2, 40.1] |
| Rationale + answer supervision | 41.79 | 84 / 201 | [35.2, 48.7] |
The 700 / 99 / 201 training, validation and test split keeps related rooms and scene families together. Both training conditions improve over the zero-shot baseline.
Rotation accuracy rises from 32.84% to 58.21% over answer supervision. Its 17 additional correct answers equal the net gain across all 201 test questions.
How to read the additional rationale gain
The extra 8.46-point gain over answer supervision remains suggestive, with an exact paired McNemar p value of 0.064 and one training seed per condition.
The paper
Abstract
Spatial intelligence, the capacity to represent and reason about physical space, underlies embodied perception and real-world visual understanding. Existing benchmarks probe it through spatial relations, perspective-taking, and mental modeling. What remains underexamined is whether a model can maintain a coherent scene representation across views, since pairwise comparison alone cannot reveal which view departs from the truth. We thus introduce Same Space or Not (SSoN), a benchmark of cross-view contradiction resolution comprising 1,000 questions from seven indoor scene sources and six categories of controlled 3D edits. Each question presents three views of an unchanged scene and one view of an edited variant. Models must identify the inconsistent view and explain the contradiction without a designated trusted reference or a marked target. Our construction process iteratively translates human review into generation constraints. Automated checks and exhaustive human validation screen out both trivial visual shortcuts and questions lacking sufficient evidence for a unique answer. Humans achieve 83.3% accuracy, compared with 39.7% for the strongest vision-language models (VLMs), which further drops to 34.9% when requiring valid visual explanation for the correct selection. We investigate prompting, camera information, and 3D reconstruction as routes for improvement. On a 100-question intervention set with a 24% baseline, prompting and 3D reconstruction remain below 50% answer accuracy, except when ground-truth camera information is supplied, which reaches 61% answer accuracy but only 48% True solve accuracy. On a held-out 201-question test set, Qwen2.5-VL-7B achieves 21.89% accuracy, 33.33% after answer-only fine-tuning, and 41.79% after fine-tuning with both rationale and answer supervision. These results identify cross-view contradiction resolution as a substantial weakness of current VLMs that can be partly addressed by targeted supervision. The 1,000 questions and fine-tuned checkpoints will be made public.
Reference
Cite this work
Coming soon.
