Show, Don't Tell

Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

“When words fall short, images give form to spatial intent.” (立象以尽意)
Xici Zhuan (Great Treatise), The Book of Changes (《周易·系辞》)
  • Xu Wang
  • Kaixiang Yao
  • Miao Pan
  • Xiaohe Zhou
  • Xuanyu Liu
  • Wenqi Zhang
  • Xuhong Zhang

Zhejiang University

Motivation

Why Evaluate Spatial Cognition in Pixels?

Spatial judgments often live in regions, marks, paths, and visual relations before they become text. The benchmark evaluates those answers without forcing every response through a text-only interface.

Motivation figure contrasting text-interface bottlenecks with visual spatial answers.
Serialization gapRegions, paths, and affordances are compressed into brittle coordinates, labels, or prose.
Metric mismatchA visual answer can be intuitively correct while remaining incompatible with text-only evaluation.
Protocolized visual answersConstrain the generated pixels, parse the visual response, and score with the original metric.

ProVisE

Protocolized Visual Evaluation

One evaluation semantics, two answer interfaces.

Protocolized visual evaluation pipeline comparing direct text answering and visual answering.
Protocolized visual evaluation converts generated pixel-space answers into benchmark-compatible predictions.

SpatialGen-Bench

A Taxonomy of Spatial Cognition

Four capability levels connect direct visual evidence to embodied spatial action.

Samples
470
Subtasks
14
Levels
4

Interactive Protocol Trace

Visual Answer Workbench

Perception · Counting

Count target objects

How many chairs are there in the image?

Counting input image.
Input image

Results

SpatialGen-Bench Leaderboard

Includes Fable 5 and other frontier models

# Model Protocol Overall Perception Understanding Reasoning Interaction

Task Matrix

Scores across all 14 subtasks

470 samples · 14 tasks

Citation

Reference

arXiv:2607.21072

@article{wang2026show,
  title = {Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text},
  author = {Wang, Xu and Yao, Kaixiang and Pan, Miao and Zhou, Xiaohe and Liu, Xuanyu and Zhang, Wenqi and Zhang, Xuhong},
  journal = {arXiv preprint arXiv:2607.21072},
  year = {2026}
}