OmniAI Group of ZJU ACES Lab
Spatial-Interactor
Learning Spatial Reasoning through Interaction
with the Observable Physical World
1 Zhejiang University2 SAP* Equal contribution
04 / Progressive Curriculum
A progressive spatial interaction curriculum
We organize supervision by the state being updated and the horizon of that update: L1 isolates external-world changes, L2 models observer-induced changes, and L3 composes successive transitions over complete trajectories.
Passive world-state transitions
The world changes; the viewpoint stays approximately fixed. Learn object motion, state changes, and single- or multi-step operations.
15,109 QA pairs
05 / Training
From local state-transition modeling to long-horizon integration
SFT learns world-state and self-state transitions from L1 and L2. OPD then uses privileged process supervision to integrate successive transitions over long trajectories.
Local State-Transition Modeling
Associate differences between observations with object operations and ego-motion, learning reusable local state-transition representations.
Long-Horizon Integration with OPD
Use training-only transition descriptions to guide reasoning over complete trajectories, alongside verifiable final-answer rewards.
06 / Main Results
Learning from interaction improves spatial reasoning
Across four backbones, Spatial-Interactor improves spatial relations, multi-view reasoning, and tasks that integrate motion evidence over longer trajectories.
Cross-benchmark generalization
Ablation studies
07 / Process and Temporal Analysis
Process dynamics and temporal evidence
Curriculum ablations, OPD optimization dynamics, and temporal diagnostics examine how answer performance and process alignment change during training.
Progressive local-transition learning
L1 and L2 contribute differently. Training them in sequence improves all four benchmarks over Base.
OPD optimization dynamics
Across four backbones, task reward rises, rollout diversity gradually narrows, and divergence from the privileged teacher decreases.
Temporal evidence and privileged-trace diagnostics
OPD retains its advantage across frame budgets. For the frozen SFT model, ordered segment descriptions help more than shuffled descriptions.
Frame-order sensitivity
Shuffling identical frames reduces Spatial-Interactor scores by 6.6 and 7.0 points, compared with 0.9 and 0.8 for the corresponding Base models.
08 / Closed-Loop Interaction
Spatial reasoning under changing observations
The same spatial updates support repeated decisions when each action changes what the model observes next.
WalkerBench
ESI-Bench
Citation
Spatial-Interactor
Learning Spatial Reasoning through Interaction with the Observable Physical World
Read the paper@misc{yao2026spatialinteractor,
title = {Spatial-Interactor: Learning Spatial Reasoning
through Interaction with the Observable
Physical World},
author = {Yao, Kaixiang and Wang, Xu and Pan, Miao
and Hu, Xiyue and Wang, Weishi
and Dahlmeier, Daniel and Chen, Jintao
and Shen, Yongliang and Zhang, Xuhong
and Zhang, Wenqi},
year = {2026},
note = {Preprint}
}






