arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BlenDAgger:用于交互式模仿学习的混合共享控制

BlenDAgger: Blended Shared Control for Interactive Imitation Learning

Cailyn Smith, Geoffrey Sun, Henny Admoni, Zackory Erickson

arXiv 2609.37599首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BlenDAgger通过共享控制混合策略与人类动作,在干预中提升模仿学习性能,实验显示自主性能显著优于HG-DAgger,并提高过渡平滑度和数据收集效率。

AI 中文摘要

机器人策略经常从人类纠正中训练,然而远程操作机器人以提供纠正既繁琐,且人类演示者并不总是最优的。我们提出了BlenDAgger(混合DAgger),一种通过使用共享控制在干预期间混合策略和演示者的动作来收集数据以训练模仿学习策略的方法。通过混合人类和策略的动作,我们旨在提高操作策略的自主性能。我们在五个操作任务中验证了我们的方法,其中两个在真实世界,三个在仿真中。与典型的人类门控纠正方法(HG-DAgger)相比,我们的方法在两个真实世界任务上实现了高出30个百分点或更多的自主性能。我们还研究了BlenDAgger允许更高自主性能的优势,发现BlenDAgger导致策略控制和人类干预之间的过渡平滑度提高57%,以及与训练数据的轨迹相似性提高14%。在一项涉及两个真实世界任务的用户研究(n=14)中,我们发现BlenDAgger导致更快的数据收集(BF=13.32),并且我们没有发现主观感知上的差异。这些结果表明,与从完全远程操作的干预中微调机器人策略的典型方法相比,混合共享控制导致更高的自主性能。

英文摘要

Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators are not always optimal. We propose Blended DAgger (BlenDAgger), an approach for collecting data to train imitation learning policies by using shared control to blend the policy's and demonstrator's actions during interventions. By blending human and policy actions, we aim to improve the autonomous performance of manipulation policies. We validate our approach across five manipulation tasks, two in the real world and three in simulation. Our approach achieves higher autonomous performance by 30 or more percentage points on two real-world tasks compared to a typical human-gated correction approach (HG-DAgger). We also investigate the advantages of BlenDAgger that allow for higher autonomous performance, finding that BlenDAgger results in 57% smoother transitions between policy control and human interventions, and 14% higher trajectory similarity to the training data. In a user study (n=14) on two real-world tasks, we find that BlenDAgger results in faster data collection (BF=13.32), and we do not find a difference in subjective perceptions. These results show that blended shared control leads to higher autonomous performance compared to typical methods for fine-tuning robot policies from fully teleoperated interventions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑