arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22521cs.RO

潜在策略引导:一种高效灵活的跨具身迁移框架

Latent Policy Steering: An Efficient and Flexible Framework for Cross-Embodiment Transfer

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Yiqi Wang, Mrinal Verghese, Jeff Schneider

AI总结:

提出潜在策略引导(LPS)框架,通过跨具身光流预训练世界模型并在目标具身上微调,在潜在空间搜索以引导策略,实现高效跨具身迁移,显著提升扩散策略和Pi0.5性能。

AI中文摘要:

学习型机器人视觉运动策略的性能在很大程度上取决于其训练数据的规模和质量,然而在现实世界中,为机器人收集高质量的示范数据仍然成本高昂。尽管大规模机器人和人类数据集日益可得,但具身差异和动作空间不匹配使得这些数据难以直接利用。跨具身迁移,即重用其他具身的经验来改进目标具身上的学习,因此对于将机器人学习扩展到超越每个机器人单独数据收集的层面至关重要。在这项工作中,我们发现,通过利用跨具身共享的信息——即世界如何响应运动的视觉动力学——并在测试时有效利用稀缺的目标具身数据,可以实现高效的迁移。所提出的框架,称为潜在策略引导(LPS),实现了一个与具身无关的预训练阶段,该阶段使用跨不同具身的光流训练一个基于图像的世界模型(WM)。得到的WM在目标具身上使用机器人动作进行微调。然后,它通过在WM的潜在空间中搜索与微调数据保持接近的计划,将基础策略引导向更好的动作。LPS是一个与策略无关的框架:它可以灵活地适应不同的策略,而无需重新训练它们。在Robomimic和真实世界评估中,LPS将扩散策略的平均性能相对提高了16%和62%,将Pi0.5提高了8%和14%,仅使用50个在未见过的目标具身上的示范。

英文摘要:

The performance of learned robot visuomotor policies depends heavily on the size and quality of their training data, yet collecting high-quality demonstrations remains costly for robots in the real world. Although large-scale robot and human datasets are increasingly available, embodiment gaps and mismatched action spaces make them difficult to leverage directly. Cross-embodiment transfer, reusing experience from other embodiments to improve learning on a target embodiment, is therefore crucial for scaling robot learning beyond per-robot data collection. In this work, we find that efficient transfer can be achieved by learning from what is shared across embodiments, the visual dynamics of how the world responds to motion, and by effectively exploiting the scarce target-embodiment data at test time. The proposed framework, called Latent Policy Steering (LPS), implements an embodiment-agnostic pretraining phase, which trains an image-based World Model (WM) with optical flow across diverse embodiments. The resulting WM is finetuned on the target embodiment with robot actions. It then steers the base policy toward better actions by searching in the WM's latent space for plans that stay close to the finetuning data. LPS is a policy-agnostic framework: it can flexibly accommodate different policies without having to retrain them. In Robomimic and real-world evaluations, LPS improves the average performance of Diffusion Policy relatively by 16% and 62%, and Pi0.5 by 8% and 14%, with only 50 demonstrations on an unseen target embodiment.

↑