arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TUCO:为仿真到现实机器人策略协同训练策划仿真演示

TUCO: Curating Simulation Demonstrations for Sim-to-Real Robot Policy Co-Training

Ning Zhu, Mengfei Zhao, Yikai Tang, Zhangyujie Sun, Peihao Li, Dongyue Ni, Jindou Jia, Jianfei Yang

arXiv 2610.05407首次发表:更新:

发表机构

Stanford University; Nanyang Technological University; AXIS Robotics; Carnegie Mellon University; Fudan University; University of California, Berkeley; Shanghai Jiao Tong University(斯坦福大学; 南洋理工大学; AXIS Robotics; 卡内基梅隆大学; 复旦大学; 加州大学伯克利分校; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对仿真到现实机器人策略协同训练,提出TUCO方法,利用影响函数分解演示贡献,统一优化轨迹效用与集合覆盖,实验证明其性能最优。

AI 中文摘要

仿真演示可以补充机器人策略协同训练中稀缺的真实世界数据。然而,利用数据策划主动选择这些演示以进行仿真到现实协同训练的价值仍未得到充分探索。现有的策划方法也缺乏统一的准则来衡量轨迹级效用和从闭环目标行为出发的集合级覆盖。为弥补这些不足,我们首次系统性地研究了仿真到现实机器人策略协同训练的数据策划,并提出了轨迹级效用与集合级覆盖优化(TUCO)。TUCO使用影响函数追踪每个源演示如何影响目标域评分轨迹。我们的关键洞见是,这些影响可以分解为对目标回报的总体贡献和跨轨迹的变异,为衡量轨迹效用和集合覆盖提供了共同的闭环基础。我们进一步提出了一种性能对齐的子集优化器,将这些度量统一到一个策划目标中,以减少冗余并选择互补的演示。在RoboMimic和OmniReset上的大量实验确立了主动仿真数据策划对仿真到现实策略协同训练的价值,并表明TUCO在单仿真器、仿真到仿真以及仿真到现实设置中均达到了最先进的性能。

英文摘要

Simulation demonstrations can supplement scarce real-world data for robot policy co-training. However, the value of using data curation to actively select these demonstrations for sim-to-real co-training remains underexplored. Existing curation methods also lack a unified criterion for measuring trajectory-level utility and set-level coverage from closed-loop target behavior. To address these gaps, we present the first systematic study of data curation for sim-to-real robot policy co-training and propose Trajectory-level Utility and set-level Coverage Optimization (TUCO). TUCO uses influence functions to trace how each source demonstration affects target-domain scoring rollouts. Our key insight is that these effects can be decomposed into an overall contribution to target return and variation across rollouts, providing a common closed-loop basis for measuring trajectory utility and set coverage. We further propose a performance-aligned subset optimizer that combines these measures in a unified curation objective to reduce redundancy and select complementary demonstrations. Extensive experiments on RoboMimic and OmniReset establish the value of active simulation data curation for sim-to-real policy co-training and show that TUCO achieves state-of-the-art performance across single-simulator, sim-to-sim, and sim-to-real settings.

Comments28 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑