OmniAI Group of ZJU ACES Lab

Spatial-Interactor

Learning Spatial Reasoning through Interaction
with the Observable Physical World

1 Zhejiang University2 SAP* Equal contribution

01 / Motivation

State-transition diagnostics reveal the missing capability

Frame shuffling exposes weak temporal-order sensitivity, while a local-to-long-horizon comparison shows that models struggle to integrate consecutive transitions over complete trajectories.

Frame shuffling barely changes Base performance. Adding transition descriptions improves long-horizon performance.

02 / State-Transition Learning

Learning physical-world state transitions through interaction

Each interaction aligns consecutive observations with the physical change that connects them. Individual records support local transition inference; ordered trajectories support trajectory-level reasoning.

Conventional spatial QA learns isolated knowledge. Spatial-Interactor learns state transitions from interactions to maintain coherent spatial states.

03 / Curriculum-Driven Construction

From interaction records to verifiable transition targets

Executed interactions and geometric trajectories provide the state changes. State, pose, and trajectory metadata determine the targets, which deterministic templates render as QA pairs.

Simulated object interactions, ego-motion, real camera trajectories, and robot trajectories feed a ground-truth-to-QA pipeline with quality filtering and visual verification.

04 / Progressive Curriculum

A progressive spatial interaction curriculum

We organize supervision by the state being updated and the horizon of that update: L1 isolates external-world changes, L2 models observer-induced changes, and L3 composes successive transitions over complete trajectories.

Passive world-state transitions

The world changes; the viewpoint stays approximately fixed. Learn object motion, state changes, and single- or multi-step operations.

15,109 QA pairs

    Task
    01 / 04
    Open / close: two observations from ProcTHOR
    SimulatedProcTHORE02

    Question

    What state change occurred between the two views?

    Reference answer

    C. An object was opened or closed.

    Composition of LSI-108K

    Full LSI-108K source distribution and taxonomy. Multi-step operations are in L1.

    05 / Training

    From local state-transition modeling to long-horizon integration

    SFT learns world-state and self-state transitions from L1 and L2. OPD then uses privileged process supervision to integrate successive transitions over long trajectories.

    Stage 1 L1 + L2

    Local State-Transition Modeling

    Associate differences between observations with object operations and ego-motion, learning reusable local state-transition representations.

    Stage 2 L3

    Long-Horizon Integration with OPD

    Use training-only transition descriptions to guide reasoning over complete trajectories, alongside verifiable final-answer rewards.

    A frozen annotator constructs ordered segment descriptions. Teacher and student evaluate the same student-generated prefix. Reasoning-only distillation and verifiable rewards jointly optimize the student.

    06 / Main Results

    Learning from interaction improves spatial reasoning

    Across four backbones, Spatial-Interactor improves spatial relations, multi-view reasoning, and tasks that integrate motion evidence over longer trajectories.

    Main benchmark results

    Main spatial reasoning benchmark results reported in the paper

    Cross-benchmark generalization

    Cross-benchmark generalization results reported in the paper

    Ablation studies

    Training ablation results reported in the paper

    07 / Process and Temporal Analysis

    Process dynamics and temporal evidence

    Curriculum ablations, OPD optimization dynamics, and temporal diagnostics examine how answer performance and process alignment change during training.

    Progressive local-transition learning

    L1 and L2 contribute differently. Training them in sequence improves all four benchmarks over Base.

    Grouped bars compare Base, L1 only, L2 only, and sequential L1 then L2 on MMSI, ViewSpatial, SAT-Real, SAT-Syn, and Average.

    OPD optimization dynamics

    Across four backbones, task reward rises, rollout diversity gradually narrows, and divergence from the privileged teacher decreases.

    Three training curves show task reward, rollout entropy, and reasoning-token process divergence for Qwen2.5-VL-3B/7B and Qwen3-VL-4B/8B.

    Temporal evidence and privileged-trace diagnostics

    OPD retains its advantage across frame budgets. For the frozen SFT model, ordered segment descriptions help more than shuffled descriptions.

    Left, Base, SFT, GRPO, and OPD route-planning scores at 16, 32, and 64 frames. Right, no trace 32.6, ordered four segments 35.4, shuffled four segments 33.8, and ordered eight segments 35.3.

    Frame-order sensitivity

    Shuffling identical frames reduces Spatial-Interactor scores by 6.6 and 7.0 points, compared with 0.9 and 0.8 for the corresponding Base models.

    Average VSTI-Bench scores fall from 48.3 to 41.7 for Spatial-Interactor-3B and from 49.5 to 42.5 for 7B after frame shuffling. A radar chart shows the task breakdown.

    08 / Closed-Loop Interaction

    Spatial reasoning under changing observations

    The same spatial updates support repeated decisions when each action changes what the model observes next.

    WalkerBench

    Table 4. WalkerBench results

    ESI-Bench

    Table 5. ESI-Bench interaction category results

    09 / Qualitative Analysis

    Spatial reasoning across views and trajectories

    A shared anchor resolves a local viewpoint change; ordered landmarks recover motion across a longer trajectory.

    Cross-view anchor correspondence

    Selected SAT-Real case with two input views and attention visualizations for Base and Spatial-Interactor.

    Long-horizon trajectory integration

    Video observations and reasoning comparison. Spatial-Interactor identifies rightward movement, while Base predicts forward movement.

    Citation

    Spatial-Interactor

    Learning Spatial Reasoning through Interaction with the Observable Physical World

    Read the paper
    BibTeX
    @misc{yao2026spatialinteractor,
      title = {Spatial-Interactor: Learning Spatial Reasoning
               through Interaction with the Observable
               Physical World},
      author = {Yao, Kaixiang and Wang, Xu and Pan, Miao
                and Hu, Xiyue and Wang, Weishi
                and Dahlmeier, Daniel and Chen, Jintao
                and Shen, Yongliang and Zhang, Xuhong
                and Zhang, Wenqi},
      year = {2026},
      note = {Preprint}
    }

    100%
    Open image