arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20114cs.AIcs.RO

DECOWAM:用于腿式移动操作的解耦全身世界-动作模型

DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

Siyuan Ma, Boshi Zhang, Yutian Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Qiaojun Yu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出DECOWAM模型,通过解耦全身动作因素提升移动操作的视觉与动作预测性能,在真实机器人数据集ARMDOG上验证了其在协调性和鲁棒性上的优势。

中文摘要 AI 辅助

移动操作要求机器人预测运动和手臂运动如何共同改变未来观测与控制。现有世界-动作模型大多为固定基座平台开发,未明确区分相机自运动与基座、手臂动作。本文提出DECOWAM,一种全身世界-动作模型,通过专用条件接口分离这些因素。DECOWAM冻结适配后的FastWAM主干网络,训练残差适配器、从特权观测中蒸馏出的动作等价未来瓶颈、对抗分离的基座与手臂潜变量,以及用于视频预测的基座速度条件。我们进一步推出ARMDOG,一种同步视频、全身状态、动作及语言的真实机器人数据集。在固定重放协议下,DECOWAM在未来视频和动作预测上均优于FastWAM,动作均方误差(MSE)降低21.7%,可训练适配参数为25.95M。在每种方法的79次闭环试验中,它在对比系统中实现了最高的全身协调性和基座位移鲁棒性,同时任务完成度与最强基准相当。这些结果表明,具身感知分解可支持在移动视角下实现参数高效的联合视觉预测与全身控制。

英文摘要

Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion from base and arm actions. Here we introduce DECOWAM, a whole-body world-action model that separates these factors through dedicated conditional interfaces. DECOWAM freezes an adapted FastWAM backbone and trains residual adapters, an action-equivalent future bottleneck distilled from privileged observations, adversarially separated base and arm latents, and base-velocity conditioning for video prediction. We further introduce ARMDOG, a real-robot dataset that synchronizes video, whole-body state and action, and language. On a fixed replay protocol, DECOWAM improved both future-video and action prediction over FastWAM, reducing action MSE by 21.7% with 25.95M trainable adaptation parameters. Across 79 closed-loop trials per method, it achieved the highest observed whole-body coordination and base-displacement robustness among the compared systems, while task completion remained comparable to the strongest baseline. These results show that embodiment-aware factorization can support parameter-efficient joint visual prediction and whole-body control under moving viewpoints.

发表机构

  • Tsinghua University(清华大学)
  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • Harbin Institute of Technology(哈尔滨工业大学)
  • Hangzhou Yunshenchu Technology Co., Ltd. (DEEP Robotics)(杭州云神初科技有限公司(深智机器人))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑