发表机构
The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对潜在动作模型跨实体迁移中背景噪声和潜在编码不一致的问题,提出动作相似性监督方法,在RoboTwin 2.0上使跨实体成功率翻倍以上。
AI 中文摘要
随着通用机器人策略从网络规模预训练中获取视觉和语言能力,演示数据的收集仍然成本高昂,并且与记录它们的机器人紧密绑定。潜在动作模型(LAMs)通过从无动作视频中学习可在不同实体间共享的潜在动作来解决这两个问题,然而在实践中,LAMs对背景视觉噪声敏感,且两台不同机器人执行相同动作可能被编码为不同的潜在变量。解决背景视觉噪声的一种方法是添加辅助损失,从潜在动作预测机器人动作,从而将潜在动作空间与特定实体的机器人动作空间关联起来。我们研究了相同标签的不同用途,即通过动作相似性监督。任意两个潜在动作之间的相似度被训练为与两条真实机器人动作序列的相似度匹配。LAM从不预测真实动作,因此潜在动作无需编码实体特定信息。我们在RoboTwin 2.0上以受控设置评估跨实体迁移:两台双臂机器人演示不相交的任务集,策略在所有演示上训练,每台机器人在仅由另一台机器人演示的任务上进行闭环评估。在固定策略架构及其超参数、数据集和评估协议的情况下,预测潜在动作而非真实动作使跨实体成功率提高了一倍以上。在给定相同真实动作的情况下,相似性监督比在LAM训练期间预测真实动作的辅助损失迁移效果更好。在末端执行器运动而非关节空间运动上计算相似度,并让损失比较两台机器人的潜在动作,是本研究中的最佳方法。
英文摘要
As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain costly to collect and tied to the robot that recorded them. Latent action models (LAMs) address both by learning latent actions from action-free videos that can be shared across embodiments, however, in practice, LAMs are sensitive to background visual noise, and the same motion from two different robots may be encoded with different latents. One solution to the background visual noise is to add an auxiliary loss predicting the robot action from the latent action, further associating the latent action space to the embodiment specific robot action space. We study a different use of the same labels, through action-similarity supervision. The similarity between any two latent actions is trained to match the similarity of the two ground-truth robot action sequences. The ground-truth actions are never predicted by the LAM, so the latent action does not need to encode embodiment specifics. We evaluate cross-embodiment transfer on RoboTwin 2.0 in a controlled setup, two bimanual robots demonstrate disjoint task sets, a policy is trained on all the demonstrations, and each robot is evaluated closed-loop on the tasks only the other demonstrated. With the policy architecture and its hyperparameters, the dataset, and the evaluation protocol fixed, predicting latent actions instead of ground-truth actions more than doubles cross-embodiment success. Given the same ground-truth actions, similarity supervision transfers better than an auxiliary loss that predicts the ground-truth action during the LAM training. Computing the similarities on end-effector motion rather than joint-space motion, and letting the loss compare latent actions across the two robots, gives the best approach of the study.