arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习超越人类所能演示的

Learning Beyond What Humans Can Demonstrate

Yuchen Song, Aditya Mittal, Unnat Jain

arXiv 2609.24996首次发表:更新:

发表机构

University of California, Irvine(加州大学尔湾分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对人类难以演示的机器人操作任务,提出GLIDE框架,通过推断失败模式并生成护栏,将数据收集成功率从0-10%提升至70-90%,并支持政策学习。

AI 中文摘要

机器人操作的行为克隆依赖于专家演示。然而,对于需要动态稳定性、精确接触时机或灵巧协调的任务,人类操作员可能难以甚至无法收集数据。我们研究这种不可行演示机制,并提出GLIDE:高效学习不可行演示的护栏框架,该框架推断任务特定的失败模式,并将其转化为可执行的数据收集和政策部署护栏。给定任务描述和条件遥操作代码,GLIDE编写护栏,利用系统状态过滤遥操作和政策命令,约束易失败动作,并基于轨迹反馈迭代改进。在三个任务中,GLIDE发现了超越领域专家硬编码护栏的新兴护栏,相较于朴素VR遥操作和领域专家硬编码护栏,提高了数据收集效率。经过改进,GLIDE将三个任务的数据收集成功率从0-10%提升至70-90%。在政策执行阶段,混合数据护栏政策在番茄盘转移、标记交接与站立、倒酒任务中分别达到70%、60%和60%的成功率。这些结果表明,当直接演示不可行时,GLIDE能够支持政策学习。项目网站:此HTTP URL。

英文摘要

Behavior cloning for robot manipulation relies on expert demonstrations. However, for tasks that require dynamic stability, precise contact timing, or dexterous coordination, human operators may find it hard or even impossible to collect data. We study this infeasible-demonstration regime and propose GLIDE: Guardrails for Learning from Infeasible Demonstrations Efficiently, a framework that infers task-specific failure modes and converts them into executable guardrails for data collection and policy deployment. Given a task description and the conditioning teleoperation code, GLIDE writes guardrails that use system states to filter teleoperation and policy commands, constrain failure-prone actions, and iteratively improve from trajectory feedback. Across three tasks, GLIDE discovers emergent guardrails that go beyond domain-expert hardcoded ones, improving data collection over naive VR teleoperation and domain-expert hardcoded guardrails. After refinement, GLIDE raises data-collection success from 0-10 percent to 70-90 percent across the three tasks. During policy execution, mixed-data guarded policies reach 70 percent, 60 percent, and 60 percent success on Tomato plate transfer, Marker handover and stand, and Wine serving tasks. These results show that GLIDE can support policy learning when direct demonstrations are infeasible. Project website: http://guardrail-policy.github.io/

CommentsAccepted to CoRL 2026. Project website: http://guardrail-policy.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑