arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mulligan:面向高效机器人学习的性能引导数据收集

Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning

Lars Ankile, Perry Dong, Rohan Bhowmik, Aneesh Muppidi, David D. Yuan, Shuran Song, Chelsea Finn

arXiv 2610.05882首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Mulligan方法,通过基于性能引导的初始状态采样和利用失败数据的价值函数,在固定预算下高效提升机器人操作策略,显著提高真实任务成功率。

AI 中文摘要

从人类示范中学习是教授机器人新任务的可靠方式,但随着策略的改进,每个额外示范带来的收益会逐渐减少。持续改进可以转而通过监督部署来实现,即操作员放置物体并在策略失败时进行干预。我们研究如何在固定预算的监督回合内,针对具有广泛物体放置范围的高精度操作任务,最大化改进效果。我们观察到,失败可能集中在初始状态的一小部分子集中,因此均匀收集会浪费操作员大量时间在策略已能处理的状态上。Mulligan将初始状态分布视为一个决策,每轮从观察到的失败和未尝试的状态开始。为了进一步提高数据效率,我们使用在所有数据(包括模仿学习丢弃的失败数据)上训练的价值函数来增强交互式模仿学习。在2,550个保留的、盲测的回合中评估的三个真实世界任务和两个模拟任务中,Mulligan在匹配的收集预算下优于均匀初始状态采样,并且结合基于价值的动作选择,HiL-IDQL+Mulligan,将最终真实任务成功率提高了10-34个百分点。在操作员干预下,人机团队完成了98%的收集回合,在策略学习期间保持生产力。视频、代码和数据可在以下网址获取:https URL。

英文摘要

Learning from human demonstrations is a reliable way to teach robots new tasks, but the gains from each additional demonstration shrink as the policy improves. Continued improvement can instead come from supervised deployment, where an operator places the objects and intervenes when the policy fails. We ask how to maximize improvement from a fixed budget of supervised episodes on high-precision manipulation tasks with wide ranges of object placements. We observe that failures can concentrate in a small subset of initial states, so uniform collection spends much of the operator's time on states the policy already handles. Mulligan makes the initial-state distribution a decision, starting each round's episodes at observed failures and untried states. To further improve data efficiency, we augment interactive imitation learning with a value function trained on all data, including failures that imitation discards. Across three real-world tasks evaluated on 2,550 held-out, blinded episodes and two simulated tasks, Mulligan outperforms uniform initial-state sampling at matched collection budgets, and combined with value-based action selection, HiL-IDQL+Mulligan, improves final real-task success by 10-34 percentage points. With operator interventions, the human-robot team completes 98% of collection episodes, remaining productive while the policy learns. Videos, code, and data are available at https://mulligan.page/.

CommentsProject page: https://mulligan.page/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑