A benchmark for spatial intelligence

Same Space or Not

Benchmarking Spatial Intelligence
Through Cross-View Contradictions

Huizi Cao1Guanyu Xu2Zhixuan Liang3Ming Yin3Jiuxiang Gu4Manling Li5Mengdi Wang3Shilong Liu3
1Carleton College2University of Pennsylvania3Princeton University4Adobe Research5Northwestern University

Four views. One contradiction.

Three images agree on one unchanged 3D scene. A fourth does not.
Can a vision-language model find it and explain why?

Try the task yourself
Four views of a room illustrate a missing basketball. A chart compares six VLMs with human performance and separates True solve from Partly guessed answers.
The same scene can look very different across cameras. SSoN asks which view cannot belong, and whether the explanation identifies a real contradiction.
83.3%Human accuracy
39.7%Best VLM accuracy
34.9%Best VLM True solve

All three headline scores use the full 1,000-question benchmark. True solve requires a correct choice and a factual, sufficient explanation.

See the challenge

Which view breaks the shared scene?

Explore one representative question for each edit type. Look across all four views, then select the inconsistent image.

These examples come from the paper. Edit labels are shown here for exploration. In the standard evaluation, models receive neither the edit label nor a marked target. Your selections stay in this page and are not submitted.

The benchmark

A contradiction must be attributable.

Finding two disagreeing images is not enough. With no trusted reference view and no marked target, the model must identify which view violates the scene supported by the others.

1,000questions
4,000rendered views
7scene sources
6edit categories
The four original Figure 3(a) donut charts show scene sources, image resolution, edit categories and object size across 1,000 questions. Their original purple, green, blue, gold, pink and orange palette is preserved.

Scroll sideways to view all four charts, or tap the figure to open it at full size.

iTHORProcTHORRoboTHOR3D-FRONTInfinigen IndoorsSAGE-10kTDW

iTHOR and ProcTHOR are reported jointly for source counts because their room configurations and components overlap in the benchmark. RoboTHOR is reported separately.

Validity-guided constraint discovery

Review of more than 10,000 development candidates turned failure cases into generation constraints. Each accepted item preserves diagnostic evidence in the intruder and at least two anchors while excluding tested 2D shortcuts.

  1. 01

    Check the scene and evidence

    Verify physical integrity, the recorded edit and clear visual evidence using the images, answer key and top-down views.

  2. 02

    Verify a unique answer

    A fresh reviewer first solves the four-image question without the key, records evidence and ambiguity, then checks for alternative answers.

  3. 03

    Reject visual shortcuts

    Exclude questions solvable through isolated appearance, object inventories or nearly aligned pixel comparisons.

View the construction overview from the paper
Original SSoN construction diagram showing scene edits, evidence requirements, validation and iterative constraint discovery.

The overview summarizes the two validation objectives. The three human checks are described above.

Full benchmark results

A wide gap remains.

Six VLMs are evaluated on the same 1,000 questions. Switch between selecting the right view and supporting that selection with valid evidence.

Download CSV ↓

Answer accuracy is the percentage of questions with the correct final choice.

Scores in percent. Each column uses all questions in that category as its denominator.
GPT-6 Astra Pro39.7092.0827.207.2364.3661.3983.17
GPT-5.6 Sol Pro24.2058.4214.406.3624.7542.5756.44
Gemini 3.5 Flash20.6046.5313.209.5419.8030.6941.58
Claude Fable 519.3039.6014.005.2024.7531.6842.57
Gemini 3.6 Flash18.7034.6516.005.7823.7629.7037.62
Grok 4.514.3022.7712.007.8023.7622.7715.84
Human83.3097.0371.6076.0195.0598.0297.03
Random choice25.0025.0025.0025.0025.0025.0025.00

Bold values mark the best VLM in each category. Human participants supplied one answer per question, yielding 833 correct choices. Human explanation scores were not measured. Seventeen non-generative vision encoders achieved 16.9% to 21.7% answer accuracy.

Rotation is particularly difficult. Even the best VLM answer accuracy in this category is 9.54%, compared with 76.01% for humans.

What SSoN reveals

Key findings

Beyond the performance gap, three findings show where models struggle across views and what helps them improve.

Six VLMs · 1,000 questions each

Correct choices can hide unsupported explanations.

404 of 1,368 correct selections lack sufficient valid evidence. True solve reveals what answer accuracy alone misses.

29.5%of correct selections
lack valid support
Explore the evidence

Accuracy decomposed

Right answer. Right reason?

True solvePartly guessed

Questions (%)

Full benchmark explanation scores. Each model is evaluated on all 1,000 questions.
ModelAnswer accuracy (%)True solve (%)Partly guessed (%)
GPT-6 Astra Pro39.734.94.8
GPT-5.6 Sol Pro24.216.97.3
Gemini 3.5 Flash20.611.69.0
Claude Fable 519.313.16.2
Gemini 3.6 Flash18.711.96.8
Grok 4.514.38.06.3
Each stacked bar shows answer accuracy. Its blue portion shows True solve. Both use all 1,000 questions per model as the denominator.
964 supported selections

A factual explanation must give sufficient evidence for the chosen view. A correct letter alone is not enough.

404 unsupported selections

Partly guessed includes fabricated changes and insufficient evidence. These cases are not all observed hallucinations.

Every correct response was reviewed by humans.

Ten Claude Opus 5 reviewer seats assisted with explanation screening. Every correct-choice response received an independent, agent-blind human judgment followed by full evidence review. Final labels are human decisions.

Download the full benchmark results ↓
100 fixed diagnostic questions

Geometric support strengthens contradiction resolution.

The reconstruction workflow improves accuracy by 23 points and True solve by 22. Supplied camera geometry achieves the highest scores.

61%accuracy with
camera geometry
Explore the evidence

Prompting, reconstruction and geometry

What helps models find the contradiction?

Answer accuracyTrue solve

Questions (%)

Overall results on the fixed 100-question diagnostic set.
ConditionAnswer accuracy (%)True solve (%)
Base2412
Comparison4023
Scene description3729
Edit taxonomy4432
Comparison + taxonomy4834
Reconstruction4734
Camera geometry6148
Darker bars show answer accuracy and lighter bars show True solve. All conditions use the same selected 100 questions. Base reuses the original responses.
+23 pointsReconstruction accuracy
over Base
+22 pointsReconstruction True solve
over Base
20 → 70%Move accuracy and True solve
with camera geometry
Camera geometry helps displacement judgments.

Adding calibration to comparison prompting raises both move scores from 20% to 70%. All 14 correct move selections have valid support.

Rotation remains difficult.

Even with camera geometry, rotation reaches 38.89% accuracy and 22.22% True solve. Reconstruction yields only two move and one rotate True solve responses.

What each condition adds

Comparison aligns landmarks, instances, expected visibility and corroborating views.

Scene description adds a written account of the 3D arrangement before deciding.

Edit taxonomy supplies all six possible edit categories without revealing the current edit.

Comparison + taxonomy combines the comparison strategy and category knowledge.

Reconstruction executes and inspects a sparse spatial 3D reconstruction without supplied geometry.

Camera geometry adds camera intrinsics, extrinsics and image dimensions to comparison prompting.

A deliberately selected diagnostic set

The set contains exactly 12 original True solve responses, 12 original Partly guessed responses and 76 wrong responses from Astra Pro. True solve and Partly guessed counts match within each edit category. Every category contributes at least 12 questions and all seven sources are represented.

Base reuses the original responses. Other conditions are new runs. These comparisons describe the selected diagnostic set rather than an average improvement over the full benchmark.

Recon uses GPT-6 Astra at the highest of six available reasoning-effort settings. Humans verify the executed spatial reconstruction, reconstructed camera views and its use in reasoning. Runs lacking execution or required evidence were repeated. The first compliant response was retained regardless of correctness. The reported gains compare the complete reconstruction workflow with Base.

Download the intervention results ↓
Qwen2.5-VL-7B · 201 held-out questions · Unseen rooms

Targeted supervision improves performance on unseen rooms.

Accuracy rises from 21.89% to 33.33% with answer supervision and 41.79% with rationale and answer supervision.

+19.90accuracy points
over zero-shot
Explore the evidence

Qwen2.5-VL-7B-Instruct

Learning to identify the contradiction

201 held-out questions

Answer accuracy (%)

Fine-tuning results on 201 held-out questions from unseen rooms.
SupervisionAccuracy (%)Correct95% Wilson confidence interval
Zero-shot21.8944 / 201[16.7, 28.1]
Answer supervision33.3367 / 201[27.2, 40.1]
Rationale + answer supervision41.7984 / 201[35.2, 48.7]
95% Wilson confidence intervals are [16.7, 28.1] for zero-shot, [27.2, 40.1] for answer supervision and [35.2, 48.7] for rationale and answer supervision.
Gains transfer to unseen rooms.

The 700 / 99 / 201 training, validation and test split keeps related rooms and scene families together. Both training conditions improve over the zero-shot baseline.

Rationale gains concentrate in rotation.

Rotation accuracy rises from 32.84% to 58.21% over answer supervision. Its 17 additional correct answers equal the net gain across all 201 test questions.

How to read the additional rationale gain

The extra 8.46-point gain over answer supervision remains suggestive, with an exact paired McNemar p value of 0.064 and one training seed per condition.

The paper

Abstract

Spatial intelligence, the capacity to represent and reason about physical space, underlies embodied perception and real-world visual understanding. Existing benchmarks probe it through spatial relations, perspective-taking, and mental modeling. What remains underexamined is whether a model can maintain a coherent scene representation across views, since pairwise comparison alone cannot reveal which view departs from the truth. We thus introduce Same Space or Not (SSoN), a benchmark of cross-view contradiction resolution comprising 1,000 questions from seven indoor scene sources and six categories of controlled 3D edits. Each question presents three views of an unchanged scene and one view of an edited variant. Models must identify the inconsistent view and explain the contradiction without a designated trusted reference or a marked target. Our construction process iteratively translates human review into generation constraints. Automated checks and exhaustive human validation screen out both trivial visual shortcuts and questions lacking sufficient evidence for a unique answer. Humans achieve 83.3% accuracy, compared with 39.7% for the strongest vision-language models (VLMs), which further drops to 34.9% when requiring valid visual explanation for the correct selection. We investigate prompting, camera information, and 3D reconstruction as routes for improvement. On a 100-question intervention set with a 24% baseline, prompting and 3D reconstruction remain below 50% answer accuracy, except when ground-truth camera information is supplied, which reaches 61% answer accuracy but only 48% True solve accuracy. On a held-out 201-question test set, Qwen2.5-VL-7B achieves 21.89% accuracy, 33.33% after answer-only fine-tuning, and 41.79% after fine-tuning with both rationale and answer supervision. These results identify cross-view contradiction resolution as a substantial weakness of current VLMs that can be partly addressed by targeted supervision. The 1,000 questions and fine-tuned checkpoints will be made public.

Reference

Cite this work

Coming soon.

View A
Enlarged question view