arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

刻意练习:有限预算下的机器人技能学习

Deliberate Practice: Learning Robot Skills under a Budget

Shivam Vats, Sudarshan Harithas, Mete Tuluhan Akbulut, Arvind Raghunathan, George Konidaris

arXiv 2608.13415首次发表:更新:

AI 中文总结

研究有限预算下机器人序列任务技能学习问题,提出刻意练习算法,通过双线性规划计算预算最优分配,经仿真与真实实验验证可提升机器人长程规划能力。

AI 中文摘要

我们研究在有限练习预算下自主学习机器人序列任务技能的问题。我们提出一种主动技能学习算法——刻意练习(Deliberate Practice,DP),该算法可计算出经证明的预算最优分配方案:在预算范围内练习可掌握且能最大化预期累积奖励的技能。DP 既估算掌握技能所需的时间,也估算该技能可解锁的任务计划的累积奖励。由于需在庞大的练习预算下对组合数量众多的技能计划进行推理,计算预算最优分配极具挑战性。我们的核心贡献是提出一种双线性规划,可通过现成求解器精确计算该分配方案。通过对长程操作任务开展的仿真与真实世界实验,我们表明所提方法能让机器人最优利用有限练习时间获取有用策略,提升长程规划能力。

英文摘要

We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks. We propose an active skill learning algorithm, \emph{Deliberate Practice (DP)}, that computes a provably \emph{budget-optimal} allocation---practicing skills that maximize expected cumulative reward while being learnable within the budget. DP estimates both the time needed to master skills and the cumulative reward of the task plans that the skills unlock. Computing a budget-optimal allocation is challenging as it requires reasoning about combinatorially many skill plans over a large practice budget. Our key contribution is a bilinear program that can compute this exactly using off-the-shelf solvers. Through simulated and real-world experiments on long-horizon manipulation tasks, we show that our approach allows robots to optimally use limited practice time to acquire useful policies and improve long-horizon planning.

Comments16 pages including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑