arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36171cs.RO

SkillWeaver:基于神经交互技能的智能体探索用于可扩展机器人数据生成

SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation

He Zhu, Lusen Zhao, Kwan Man Cheng, Su Li, Katerina Fragkiadaki

首次发表
浏览论文内容

中文总结 AI 辅助

提出SkillWeaver框架,通过VLM智能体探索神经交互技能并利用验证器引导的树搜索,自主生成39.1K演示数据,显著提升机器人策略的泛化与迁移能力。

中文摘要 AI 辅助

大规模演示数据推动了机器人学习的空前进步,然而通过遥操作收集机器人数据成本高昂,且难以扩展到多样化的环境和长时程任务。仿真提供了一种可扩展的替代方案,但现有的数据生成流程通常依赖于开环控制器、脚本化技能序列或特定任务的程序。我们提出了SkillWeaver,一个智能体框架,通过探索神经交互技能(NIS)自主生成机器人经验:NIS是可复用、参数化的闭环策略,将学习到的物理交互能力暴露给推理智能体。给定任务和仿真环境,VLM智能体推理下一步行动,调用并参数化NIS与环境交互,观察其结果,并生成验证、反思和记忆以指导后续探索。我们将NIS实例化为用于闭环、接触丰富操作的强化学习策略,并将探索组织为验证器引导的树搜索,使智能体能够发现成功的长时间行为,而无需依赖预定的执行流程。SkillWeaver自主扩展到39.1K个演示,覆盖14.1K个场景,我们将其蒸馏为视觉运动策略。在仿真基准和真实世界操作中,基于SkillWeaver生成的经验进行训练,显著提高了对新颖物体、空间配置、任务和环境的泛化能力,并实现了零样本和少样本的仿真到仿真及仿真到现实的迁移。我们的结果表明,基于神经交互技能的智能体探索是机器人数据生成的一种可扩展替代方案。

英文摘要

Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines often rely on open-loop controllers, scripted skill sequences, or task-specific programs. We introduce SkillWeaver, an agentic framework that autonomously generates robot experience by exploring over Neural Interaction Skills (NIS): reusable, parameterized, closed-loop policies that expose learned physical interaction capabilities to a reasoning agent. Given a task and a simulated environment, a VLM agent reasons about what to do next, invokes and parameterizes NIS to interact with the environment, observes their outcomes, and generates verification, reflection, and memory to guide subsequent exploration. We instantiate NIS as reinforcement-learned policies for closed-loop, contact-rich manipulation and organize exploration as verifier-guided tree search, enabling the agent to discover successful long-horizon behaviors without relying on predetermined execution pipelines. SkillWeaver scales autonomously to 39.1K demonstrations across 14.1K scenes, which we distill into visuomotor policies. Across simulation benchmarks and real-world manipulation, training on SkillWeaver-generated experience substantially improves generalization to novel objects, spatial configurations, tasks, and environments, and enables zero- and few-shot sim-to-sim and sim-to-real transfer. Our results suggest agentic exploration over neural interaction skills as a scalable alternative for robot data generation.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑