arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23038cs.AIcs.LG

Spatial-Interactor:通过与可观测物理世界的交互学习空间推理

Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World

  • Zhejiang University(浙江大学)
  • SAP

机构由 AI 辅助整理,请以论文原文为准。

Kaixiang Yao, Xu Wang, Miao Pan, Hu Xiyue, Weishi Wang, Daniel Dahlmeier, Jintao Chen, Yongliang Shen, Xuhong Zhang, Wenqi Zhang

AI总结:

针对VLM在动态环境中空间推理能力不足的问题,提出Spatial-Interactor框架,通过三级课程和两阶段训练(SFT与OPD)利用交互轨迹学习状态转换,构建LSI-108K数据集,在多个基准上取得一致提升。

AI中文摘要:

空间推理对于视觉语言模型(VLMs)理解和作用于物理世界至关重要。在动态环境中进行推理要求VLMs感知由物体运动和视角变化引起的局部状态转换,并在长轨迹上整合这些转换以维持更新的空间状态,然而现有VLMs在这两种能力上均存在局限。当前的空间训练主要聚焦于关于物体属性和空间关系的静态问题,为状态转换提供的直接监督有限;相比之下,交互轨迹自然地将先前的观察、一个动作和随后的观察连接起来,为局部状态转换提供直接监督,而完整轨迹则揭示了连续转换之间的依赖关系。因此,我们引入了Spatial-Interactor,一个通过交互训练VLMs建模物理世界状态转换的框架,将该学习过程组织为三级课程,涵盖L1被动世界状态转换、L2主动自身状态转换和L3长时程交互轨迹。相应地,我们从模拟和真实交互轨迹构建了从空间交互学习数据集(LSI-108K),其任务与每个级别的目标对齐。我们的两阶段训练策略将监督微调(SFT)应用于L1和L2以进行局部转换建模,然后在线策略蒸馏(OPD)利用特权自蒸馏:一个给定片段级转换描述的教师分支监督学生的在线策略思维链(CoT),帮助学生学会在L3长轨迹上整合连续转换。在多个VLMs和空间基准上的实验显示,在局部转换建模和长时程整合方面取得了一致的提升。

英文摘要:

Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an updated spatial state, yet existing VLMs remain limited in both capabilities. Current spatial training primarily focuses on static questions about object attributes and spatial relations, providing limited direct supervision for state transitions; in contrast, interaction trajectories naturally connect a preceding observation, an action, and a subsequent observation, offering direct supervision for local state transitions, while complete trajectories reveal dependencies among consecutive transitions. We therefore introduce Spatial-Interactor, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories. Accordingly, we construct the Learning from Spatial Interaction dataset (LSI-108K) from simulated and real interaction trajectories, with tasks aligned with the objective of each level. Our two-stage training strategy applies Supervised Fine-Tuning (SFT) to L1 and L2 for local transition modeling, and On-Policy Distillation (OPD) then uses privileged self-distillation: a teacher branch given segment-level transition descriptions supervises the student's on-policy CoT, helping the student learn to integrate consecutive transitions over L3 long trajectories. Experiments across multiple VLMs and spatial benchmarks show consistent gains in local transition modeling and long-horizon integration.

↑