技能空间射击:用于自主机器人策略改进
Skill-Space Shooting for Autonomous Robot Policy Improvement
浏览论文内容
中文总结 AI 辅助
提出技能空间射击方法,利用基础模型指导通过可重用技能探索修正,将成功试验转化为策略改进,实现自主机器人在任务内和跨任务的可扩展、可泛化策略提升。
中文摘要 AI 辅助
部署在物理世界中的机器人必须能够在遇到新情况和失败时,超越其初始训练进行改进。为了使这种改进能够跨任务扩展,必须有效利用经验,而不需要人类对每个修正进行示范。最近的智能体系统提供了一种减少对人类努力依赖的方法,通过使用基础模型自主组合学习到的行为来完成任务。然而,以这种方式完成任务本身并不能教会任务策略克服自身的失败;这需要将这些行为转化为策略的可学习修正。我们的见解是,许多这样的修正是熟悉的短行为,即技能:它们跨任务重复出现,并描述了基础模型可以从场景中推理出的动作。我们引入了技能空间射击,它利用基础模型指导,通过这些可重用的技能探索修正,并将成功的试验转化为策略改进。真实世界实验表明,自主行动的策略能够反复改进,同时技能也可以共享,以减少改进新任务所需的教导。通过使可重用技能成为修正监督的来源,技能空间射击实现了任务内和跨任务的可扩展且可泛化的策略改进。更多结果和视频见此https URL。
英文摘要
Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experience without requiring human demonstration of each correction. Recent agentic systems offer a way to reduce this reliance on human effort by using foundation models to autonomously compose learned behaviors to complete tasks. Yet completing tasks this way does not itself teach a task policy to overcome its own failures; that requires turning these behaviors into learnable corrections for the policy. Our insight is that many such corrections are familiar short behaviors, or skills: they recur across tasks and describe actions that foundation models can reason about from a scene. We introduce skill-space shooting, which uses foundation model guidance to explore corrections through these reusable skills and turn successful trials into policy improvement. Real-world experiments show repeated improvement in policies acting autonomously, while skills can also be shared to reduce the teaching needed to improve on new tasks. By making reusable skills a source of corrective supervision, skill-space shooting enables scalable and generalizable policy improvement within and across tasks. Additional results and videos at https://skill-space-shooting.github.io.
发表机构
- Tsinghua University(清华大学)
- UC Berkeley(加州大学伯克利分校)
- Shanghai Qi Zhi Institute(上海期智研究院)
机构由 AI 辅助整理,请以论文原文为准。