Motivation
Why Evaluate Spatial Cognition in Pixels?
Spatial judgments often live in regions, marks, paths, and visual relations before they become text. The benchmark evaluates those answers without forcing every response through a text-only interface.
Serialization gapRegions, paths, and affordances are compressed into brittle coordinates, labels, or prose.
Metric mismatchA visual answer can be intuitively correct while remaining incompatible with text-only evaluation.
Protocolized visual answersConstrain the generated pixels, parse the visual response, and score with the original metric.
ProVisE
Protocolized Visual Evaluation
One evaluation semantics, two answer interfaces.
SpatialGen-Bench
A Taxonomy of Spatial Cognition
Four capability levels connect direct visual evidence to embodied spatial action.
- Samples
- 470
- Subtasks
- 14
- Levels
- 4
Interactive Protocol Trace
Visual Answer Workbench
Perception · Counting
Count target objects
How many chairs are there in the image?
Results
SpatialGen-Bench Leaderboard
| # | Model | Protocol | Overall | Perception | Understanding | Reasoning | Interaction |
|---|
Task Matrix
Scores across all 14 subtasks
@article{wang2026show,
title = {Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text},
author = {Wang, Xu and Yao, Kaixiang and Pan, Miao and Zhou, Xiaohe and Liu, Xuanyu and Zhang, Wenqi and Zhang, Xuhong},
journal = {arXiv preprint arXiv:2607.21072},
year = {2026}
}