基于行为对齐表征的跨具身迁移
Cross-Embodiment Transfer via Behavior-Aligned Representations
浏览论文内容
中文总结 AI 辅助
本研究提出在视觉-语言-动作模型中采用行为对齐表征,开发仿真基准验证其可提升跨具身迁移,使真实机器人任务完成进度提升28%。
中文摘要 AI 辅助
近期,机器人操作的大规模模仿学习进展得益于利用各类机器人具身形态的数据集,但实现显著的跨具身迁移仍具挑战。本研究探讨在视觉-语言-动作(VLA)模型中采用行为对齐表征(如目标边界框、语言指令动作、机器人运动的末端执行器轨迹)对促进跨具身迁移的作用,假设这类表征具备跨具身不变性且可预测机器人动作,能统一大规模跨具身数据以增强迁移效果。为验证假设,开发了基于仿真的基准,用于评估多样跨具身数据向新具身形态的迁移效果,对比不同表征及其融入方式,发现末端执行器轨迹对迁移尤其有益,表征在更大先验数据集下更有用,还可利用无动作数据;同时证明其能增强仿真到现实的跨具身迁移,使预训练于仿真数据的真实机器人策略任务完成进度提升28%,评估视频可在指定网站查看。
英文摘要
Recent progress in large-scale imitation learning for robot manipulation has been driven by leveraging datasets across a wide range of robot embodiments. However, achieving significant cross-embodiment transfer is often still challenging. In this work, we study the role of using behavior-aligned representations (e.g., object bounding boxes, language motions, end-effector traces of robot motion) in vision-language-action (VLA) models to promote cross-embodiment transfer. We hypothesize that by possessing invariances across embodiments while being predictive of robot actions, these representations can help unify large-scale cross-embodiment data to enhance transfer. To assess our hypothesis, we develop a simulation-based benchmark designed to assess transfer with diverse cross-embodiment data to new embodiments. Using this benchmark, we compare different representations and ways of incorporating them. We identify that end-effector traces can be particularly beneficial for transfer, representations are generally more useful with larger prior datasets, and can be used to benefit from action-free data. We also demonstrate that they can enhance sim-to-real cross-embodiment transfer, improving task completion progress of real robot policies pre-trained on simulation data by 28%. We provide videos of our evaluations at our website: https://ajaysridhar.com/barx/.
发表机构
- Stanford University(斯坦福大学)
- Toyota Research Institute(丰田研究所)
机构由 AI 辅助整理,请以论文原文为准。