发表机构
ETH Zurich; Technical University of Munich; University of British Columbia; University of California, Berkeley; Google; Microsoft(苏黎世联邦理工学院; 慕尼黑工业大学; 不列颠哥伦比亚大学; 加州大学伯克利分校; 谷歌; 微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ChronoGraph,一种功能性4D场景图,将行动与状态变化关联,通过自动构建的基准和两阶段训练的VLM,统一交互理解与具身规划,实验验证了其有效性和迁移能力。
AI 中文摘要
具身智能体必须确定在哪里行动,预测由此产生的场景变化,并解释观察到的结果以指导后续行动。这需要将4D交互理解(解释过去的行动如何改变场景)与空间具身规划(确定如何以及在哪里朝着目标行动,并预测由此产生的场景变化)联系起来。我们引入了ChronoGraph,一种功能性4D场景图,它将可供性部件上的行动与语义和几何状态变化联系起来。通过以相同形式表示观察到的和预期的转变,它为理解和规划提供了共同的基础。我们通过一个自动数据引擎构建了ChronoGraphBench,该引擎将人类交互视频和模拟机器人轨迹转换为图标注的问题,用于训练和评估视觉语言模型(VLMs)在这两项任务上的表现。利用这些标注,我们通过两阶段适配预训练的视觉语言模型来训练ChronoGraphVLM。图作为思维链的监督微调教会模型在回答问题之前重建观察到的转变并预测未来的转变作为图轨迹。随后的联合4D图强化学习直接奖励图属性和答案正确性。跨模型规模的实验表明,相对于相应的预训练基线有改进,并且对VLM4D具有零样本迁移能力。真实世界演示进一步表明,基于图的规划和可供性具身能够通过现有的机器人技能支持移动操作,而无需额外的微调。
英文摘要
Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.