arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

专家示范轮廓控制:从慢速专家与快速试玩中实现超越示范速度的规划

Expert-Play Contouring Control: Faster-than-Demonstration Planning from Slow Expert and Fast Play

Seunghoon Cho, Wonsuhk Jung, Sundhar Vinodh Sangeetha, Shreyas Kousik

arXiv 2609.22798首次发表:更新:

发表机构

Seoul National University; Georgia Institute of Technology(首尔大学; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对模仿学习无法超越示范速度的问题,提出专家示范轮廓控制(EPCC),结合慢速专家示范与快速试玩数据,利用世界模型优化动作,在视觉运动操作任务上实现2.0倍于基线的吞吐量。

AI 中文摘要

专家示范通常规定了机器人应该做什么,但并未规定其能多快地完成。模仿学习(IL)继承了示范的时间安排,而直接加速学习到的运动可能会失败,因为更快的执行会改变机器人-物体的动力学。我们将超越示范速度的执行作为一个动力学感知的控制问题来研究,并提出了专家示范轮廓控制(EPCC),该方法将慢速专家示范与快速的、非专家的试玩相结合。专家示范用于训练一个潜在轨迹生成器,其预测被重新参数化为一个与时间无关的成功任务进展轮廓,而试玩数据则用于训练一个关于快速动作结果的世界模型(WM)。在部署时,我们提出的规划器利用世界模型来优化动作,这些动作在最大化沿专家轮廓进展的同时,惩罚偏离预期任务演化的行为。在三个视觉运动操作任务上取平均,EPCC 实现了模仿学习基线 2.0 倍的吞吐量,包括在一个具有交互敏感物体动力学的任务上,相对于最强加速基线 2.2 倍的吞吐量。我们的分析表明,增益集中在更快的执行改变机器人-物体演化的地方。综合来看,我们的结果凸显了一个简单而有效的原则:示范提供了任务意图,而试玩数据则提供了更快执行该意图所需的动力学覆盖。

英文摘要

Expert demonstrations often specify what a robot should do, but not how fast it can do it. Imitation Learning (IL) inherits demonstration timing, while directly accelerating the learned motion can fail when faster execution changes the robot-object dynamics. We study faster-than-demonstration execution as a dynamics-aware control problem and introduce Expert-Play Contouring Control (EPCC), which combines slow expert demonstrations with fast, non-expert play. Expert demonstrations train a latent trajectory generator whose predictions are reparameterized into a time-independent contour of successful task progression, while play trains a world model (WM) of fast-action outcomes. At deployment, our proposed planner uses the WM to optimize actions that makes maximize progress along the expert-derived contour while penalizing deviation from the intended task evolution. Averaged across three visuomotor manipulation tasks, EPCC achieves a $2.0\times$ the throughput of the IL baseline, including $2.2\times$ that of the throughput of the strongest acceleration baseline on a task with interaction-sensitive object dynamics. Our analysis shows that the gains concentrate where faster execution changes robot-object evolution. Together, our results highlight a simple yet effective principle: demonstrations provide task intent, while play data provides the dynamic coverage needed to execute that intent faster.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑