arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AutodidactWAM:从生成视频到机器人动作的跨模态自蒸馏

AutodidactWAM: Cross-Modal Self-Distillation from Generated Video to Robot Actions

Sergei Kurchev, Iaroslav Kolomiets, Miguel Altamirano Cabrera, Artem Lykov, Dzmitry Tsetserukou

arXiv 2610.08119首次发表:更新:

发表机构

Skolkovo Institute of Science and Technology (Skoltech); MWS R&D Center(斯科尔科沃科学技术学院; MWS研发中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对世界动作模型在跨本体适配中的视频-动作不对称问题,提出AutodidactWAM,利用生成视频经手部姿态估计与逆运动学恢复动作,结合DPO+SFT+DTW混合目标微调,显著提升真实机器人任务成功率。

AI 中文摘要

世界动作模型(WAMs),如Cosmos 3,能够根据观察和指令联合生成未来视频和机器人动作。将这样一个模型通过轻量级LoRA微调适配到一台未见过的机器人——配备五指BrainCo灵巧手的Unitree G1人形机器人上,暴露出视频-动作不对称性:视频渲染出合理的任务执行过程,而共同生成的动作却系统性偏离目标。我们在三个累积阶段评估了闭环真实机器人试验:预抓取、抓取和拾放。原生动作的成功率分别仅为约17%、10%和7%,且在未见过的物体上表现更差。我们提出AutodidactWAM,一种无需遥操作训练的手部姿态估计器,随后进行逆运动学计算,作用于模型生成的视频以恢复动作估计。与原生预测配对后,这些恢复的动作提供了仅微调动作相关层的优选目标,而生成的视频则采用教师强制方式。我们将监督式重标注与基于整流流的Diffusion-DPO适配进行比较。经过一次性本体适配后,自蒸馏无需额外的任务特定遥操作。恢复动作门控在预抓取、抓取和拾放阶段分别达到约75%、47%和42%的成功率,而原生动作仅为17%、10%和7%。结合偏好监督、监督目标拟合和笛卡尔轨迹锚定的混合目标(DPO+SFT+DTW)表现最佳:在训练物体Oreo上,预抓取成功率达90%,完整任务成功率达20%;在未见物体上分别达到80%和30%。纯Flow-DPO尽管验证偏好准确率达1.000,但成功率为0%,表明训练目标的组合而非对比目标本身驱动了观察到的性能提升。

英文摘要

World-action models (WAMs) such as Cosmos 3 jointly generate future video and robot actions from an observation and instruction. Adapting one such model with a lightweight LoRA fine-tune to a previously unseen robot, a Unitree G1 humanoid with five-fingered BrainCo hands, exposes a video-action asymmetry: the video renders plausible task executions, while the co-generated action is systematically mis-targeted. We evaluate closed-loop real-robot trials at three cumulative stages: pre-grasp, grasp, and pick-and-place. The native action succeeds only approximately 17%, 10%, and 7% of the time, respectively, and performs worse on held-out objects. We propose AutodidactWAM, a hand-pose estimator trained without teleoperation, followed by inverse kinematics, that runs on the model's generated video to recover action estimates. Paired with the native prediction, these recovered actions provide preferred targets for fine-tuning only the action-related layers, while the generated video is teacher-forced. We compare supervised relabeling with a rectified-flow adaptation of Diffusion-DPO. After one-time embodiment adaptation, self-distillation requires no additional task-specific teleoperation. The recovered-action gate reaches approximately 75%, 47%, and 42% pre-grasp, grasp, and pick-and-place success, compared with 17%, 10%, and 7% for the native action. A hybrid objective combining preference supervision, supervised target fitting, and Cartesian trajectory anchoring (DPO+SFT+DTW) performs best: on Oreo, the training object, it reaches 90% pre-grasp and 20% full-task success; on a held-out object, it reaches 80% and 30%. Plain Flow-DPO reaches 0% success despite 1.000 validation preference accuracy, indicating that the combination of training objectives, rather than the contrastive objective alone, drives the observed gains.

Comments8 pages,3 figures, 1 table, ICRA 2027 submitted

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑