arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考用于组合式与上下文内机器人操作的世界-动作模型

Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation

Shukai Gong, Xuanran Zhai, Yintianrun Zhang, Ruopeng Cui, Ye Huang, Yiyang Fu, Dexuan Lyu, Chaojie Li, Xinyi Song, Peiwen Lin, Chuang Wang, Mingyuan Jia, Yufan Deng, Jiaxin Fang, Bo Liang, Jiaxin Li, Yuxiang Gao, Hao Liu, Daquan Zhou

arXiv 2610.02368首次发表:更新:

发表机构

Peking University; AgiBot; CocoMatrix(北京大学; AgiBot(智元机器人); CocoMatrix(可可矩阵))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ViGAR分层框架,将组合式机器人操作分解为视觉子目标规划与执行,共享世界模型并支持上下文内学习,在RoboTwin基准上超越基线12.86个百分点,真实实验验证有效。

AI 中文摘要

长时程组合式操作对于真实世界机器人部署变得越来越重要,其中单个任务涉及多个协调的子任务。现有的世界-动作模型(WAMs)联合预测短时程视觉未来和动作,但通常缺乏显式的子任务级推理。我们提出视觉目标条件动作推理(ViGAR),一种分层框架,将操作分解为视觉子目标规划器和子目标执行器。给定当前观察和全局指令,子目标规划器预测下一子任务的视觉子目标。子目标执行器随后在预测的子目标条件下联合生成未来视觉轨迹和动作。两个组件共享预训练的世界模型表示,使任务级规划和动作生成能够受益于共同的物理知识。此外,我们的框架自然支持上下文内学习:使用全局目标图像作为上下文可以诱导不同的子任务分解和行为,而无需参数更新。在RoboTwin Clean2Random基准上,ViGAR在Clean和Random设置下分别达到82.00%和67.02%的成功率,平均成功率超过最强基线12.86个百分点。在五个组合式任务和两个上下文内学习任务上的真实世界机器人实验进一步证实了ViGAR的有效性。

英文摘要

Long-horizon compositional manipulation has become increasingly important for real-world robot deployment, where a single task involves multiple coordinated subtasks. Existing world-action models (WAMs) jointly predict short-horizon visual futures and actions, but typically lack explicit subtask-level reasoning. We propose Visual Goal-conditioned Action Reasoning (ViGAR), a hierarchical framework that factorizes manipulation into a visual subgoal planner and a subgoal executor. Given the current observation and global instruction, the subgoal planner predicts a visual subgoal for the next subtask. The subgoal executor then jointly generates future visual trajectories and actions conditioned on the predicted subgoal. Both components share a pretrained world-model representation, enabling task-level planning and action generation to benefit from common physical knowledge. Moreover, our framework naturally supports in-context learning: using a global goal image as context can induce different subtask decompositions and behaviors without parameter updates. On the RoboTwin Clean2Random benchmark, ViGAR achieves 82.00% and 67.02% success rates under the Clean and Random settings, respectively, surpassing the strongest baseline by 12.86 percentage points in average success rate. Real-world robot experiments on five compositional and two in-context learning tasks further confirm the effectiveness of ViGAR.

Comments19 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑