arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08123cs.ROcs.AIcs.LG

DISEIL:面向样本高效模仿学习的演示蒸馏

DISEIL: Demonstration Distillation for Sample-Efficient Imitation Learning

Suyog Khanal, Arun Kumar A, Santu Rana

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出DISEIL,通过有意选择失败模式和演示起点,用视觉-语言模型生成演示请求,在模拟任务中显著提升样本效率,实现高成功率。

中文摘要 AI 辅助

一个能够从少量演示中学会新任务的机器人,必须自行判断自己尚不能完成什么,然后精确地请求相应的帮助。交互式模仿学习朝这个方向迈出了一步,它让策略自行练习,并在出错时调用专家。现有方法决定何时中断学习者。另外两个决策则留给了恰好触发中断的那个回合:纠正哪个失败,以及演示应从何处开始。本文首次尝试有意地做出这两个决策。DISEIL(面向样本高效模仿学习的演示蒸馏)在策略首次变得不可靠的步骤处标记每个失败回合,用几何描述符表示该时刻,并将失败归类为反复出现的失败模式。一个视觉-语言模型和一个语言模型读取所选模式,并撰写下一个演示的请求,同时一个任务约束存储库会在花费任何专家时间之前检查该请求是否可行。没有任何模型产生机器人动作。在5个模拟任务中,在状态和图像观测下,仅改变专家被请求的内容,在所有10种设置中获得了最高的平均留出成功率,其中1种设置并列,且在我们测试的最小预算下差距最大。范围很窄:一次只进行一轮练习,在模拟环境中,专家大多是脚本化的。长期目标是让学习者也能跟踪其演示集已覆盖的内容,并根据每个请求给人类教师带来的努力程度,按比例请求缺失的行为。

英文摘要

A robot that can be taught a new task from a handful of demonstrations has to work out for itself what it still cannot do, and then ask for exactly that. Interactive imitation learning takes a step in that direction by letting a policy practice on its own and calling an expert when it goes wrong. Existing methods decide when to interrupt the learner. A further 2 decisions are left to whichever episode happened to trigger the interruption: which failure to correct, and where the demonstration should start. This paper is a first attempt at making both of them deliberately. DISEIL (Demonstration dIstillation for Sample-Efficient Imitation Learning) marks each failed episode at the step where the policy first becomes unreliable, represents that moment with a geometric descriptor, and groups the failures into recurring failure modes. A vision-language model and a language model read the selected mode and write a request for the next demonstration, and a store of task constraints checks that the request can be carried out before any expert time is spent. No model produces a robot action. Across 5 simulated tasks under state and image observations, changing only what the expert is asked for gives the highest mean held-out success rate in all 10 settings, with a tie in 1, and the margin is widest at the smallest budget we tested. The scope is narrow: a single round of practice at a time, in simulation, with experts that are mostly scripted. The longer-term aim is a learner that also tracks what its demonstration set already covers, and that asks a human teacher for the missing behavior in proportion to the effort each request costs them.

发表机构

  • Deakin Applied Artificial Intelligence Initiative(迪肯应用人工智能倡议)

机构由 AI 辅助整理,请以论文原文为准。

↑