arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Reward-DAgger:基于通用进度奖励模型的机器人门控交互式模仿学习

Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models

Ryan Li, Yigit Korkmaz, Erdem Bıyık

arXiv 2610.04054首次发表:更新:

发表机构

University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Reward-DAgger提出一种机器人门控交互式模仿学习框架,利用通用进度奖励模型检测失败并触发人工干预,无需任务特定调参,在八个模拟和真实任务中提升自主成功率并优于现有基线。

AI 中文摘要

机器人学习的最新进展使得通用控制策略能够完成广泛的任务。然而,当部署在未见过的环境中时,其性能会下降,因此检测失败并教授恢复行为变得至关重要。现有的运行时监控方法通常需要针对任务和策略进行特定训练或超参数调整,这限制了跨任务部署,并在迭代策略更新过程中引入了额外开销。我们提出了Reward-DAgger,一种机器人门控的交互式模仿学习框架,它利用来自通用奖励模型的密集进度信号来确定何时需要人工干预。我们的方法对底层策略架构不可知,无需访问策略内部,并且可以在不重新调整门控机制的情况下跨任务应用。结果表明,Reward-DAgger在失败检测的准确性与延迟权衡方面优于现有的运行时监控基线。在八个模拟和真实世界任务中,Reward-DAgger在整个交互学习过程中持续提高下游策略的自主成功率,并在大多数设置中实现了较高的人力回报,优于基线。重要的是,相同的门控配置在任务间使用,无需针对特定任务的超参数调整,展示了跨任务、环境和策略架构的迁移能力。代码和视频可在以下网址获取:此https URL。

英文摘要

Recent advances in robot learning have enabled generalist control policies capable of completing a wide range of tasks. However, their performance degrades when deployed in unseen environments, making it critical to detect failures and teach recovery behaviors. Existing runtime monitoring methods often require task- and policy-specific training or hyperparameter tuning, limiting cross-task deployment and introducing additional overhead during iterative policy updates. We present Reward-DAgger, a robot-gated interactive imitation learning framework that uses dense progress signals from a general-purpose reward model to determine when human intervention is needed. Our approach is agnostic to the underlying policy architecture, requires no access to policy internals, and can be applied across tasks without retuning the gating mechanism. Our results show that Reward-DAgger achieves a better failure-detection accuracy-latency tradeoff than existing runtime monitoring baselines. Across eight simulated and real-world tasks, Reward-DAgger consistently improves the downstream policy's autonomous success rate throughout interactive learning and achieves strong return on human effort, outperforming the baselines in most settings. Importantly, the same gating configuration is used across tasks without task-specific hyperparameter tuning, demonstrating transfer across tasks, environments, and policy architectures. Code and videos are available at https://liralab.usc.edu/reward-dagger.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑