arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探索、执行、进化:具身智能体的技能获取与复用循环

Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents

Sicheng Xie, Yitong Chen, Haidong Cao, Shunlin Lu, Zuxuan Wu, Yu-Gang Jiang

arXiv 2609.37810首次发表:更新:

发表机构

Fudan University; Shanghai Innovation Institute; NeoteAI(复旦大学; 上海创新研究院; NeoteAI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对具身智能体泛化难、执行成本高的问题,提出RoboSkill框架,通过探索-执行-进化循环获取并复用技能,结合触觉反馈与代码增强,在LIBERO-10和真实机器人上显著提升成功率并降低运行时间。

AI 中文摘要

视觉-语言-动作模型和世界-动作模型在机器人领域展示了令人印象深刻的能力,但泛化到未见过的任务仍然具有挑战性。最近,通用多模态智能体在零样本机器人任务解决方面显示出巨大潜力。然而,它们常常因从头推理和探索物理世界而产生高执行成本。为降低这些成本,我们引入了RoboSkill,一个通过探索、执行、进化循环连接技能获取与复用的框架。在该循环中,智能体进行探索以收集任务相关信息,执行任务时适应反馈,并根据执行记录进化其技能库。然后,它复用这些技能以指导下一周期的探索和执行,从而闭合循环。为提高循环效率,我们将视觉与触觉反馈相结合,以减少物理交互中的不确定性。我们进一步用可复用代码增强文本指导,以减少技能复用时的推理开销。在LIBERO-10上,RoboSkill将四种智能体的首次尝试成功率提高了12.5至25.0个百分点,并将平均运行时间减少了7.6%至72.4%。在真实机器人上,它提高了8.3个百分点的成功率,并将成功试验的平均运行时间至少减少了14.4%。

英文摘要

Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by reasoning and exploring the physical world from scratch. To reduce these costs, we introduce RoboSkill, a framework that connects skill acquisition and reuse through an Explore, Execute, Evolve loop. Within this loop, the agent explores to gather task-relevant information, executes tasks while adapting to feedback, and evolves its skill library based on execution records. It then reuses these skills to guide exploration and execution in the next cycle, closing the loop. To improve loop efficiency, we complement vision with tactile feedback to reduce uncertainty during physical interaction. We further augment textual guidance with reusable code to reduce reasoning overhead during skill reuse. On LIBERO-10, RoboSkill improves first-episode success rates by 12.5--25.0 percentage points and reduces average runtime by 7.6--72.4% across four agents. On real robots, it improves success rates by 8.3 percentage points and reduces average runtime for successful trials by at least 14.4%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑