arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EmbodiedSkills:用于编排、训练和部署VLA智能体的统一框架

EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents

Wei Wang, Wenqiao Zhang, Yutong Lin, Yuqian Yuan, Tianwei Lin, Jinhao Mao, Zhenxuan Fan, Mingjian Gao, Yang Dai, Wentong Li, Zheqi Lv, Zheng Dong, Yingjie Niu, Jiaqi Zhu, Jun Xiao, Chao Li, Yueting Zhuang

arXiv 2609.01281首次发表:更新:

发表机构

College of Computer Science and Technology, Zhejiang University; Nanjing University of Aeronautics and Astronautics; Cornell University; Universal Ubiquitous AI Co., Ltd.; Hangzhou DEEP Robotics Technology Co., Ltd.; National University of Singapore(浙江大学计算机科学与技术学院; 南京航空航天大学; 康奈尔大学; Universal Ubiquitous AI有限公司; 杭州深度机器人科技有限公司; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出EmbodiedSkills统一框架,通过可执行技能接口整合VLA智能体的感知、规划等环节,经实例化验证其在RoboTwin 2.0、LIBERO等任务上的执行性能,为构建闭环具身系统提供支持。

AI 中文摘要

视觉-语言-动作(Vision-language-action,VLA)模型将视觉观测和语言指令直接映射到机器人动作,但长时程任务所需的远不止动作预测。随着物理状态演变,智能体必须协调感知、规划、执行、进度验证与恢复环节。仅靠动作预测或模型生成的技能决策,无法保证所提议操作在当前状态下有效,也无法保证其结果能得到验证。我们提出EmbodiedSkills,这是一个将每个技能决策视为执行提议的统一框架:运行时在执行前检查其前提条件,执行后验证结果。共享的可执行技能接口将高层技能选择、有界低层VLA执行及动作后验证连接在单个智能体循环内。由于该接口保持固定,低层VLA策略可被替换或适配,无需修改智能体循环。该接口还将规划、执行、验证与恢复事件记录为结构化轨迹,这些轨迹为各组件提供监督,并在有交互反馈时支持可选的在线适配。我们在RoboTwin 2.0和LIBERO上用Qwen3-VL与OpenPI/pi0.5实例化EmbodiedSkills。经任务适配的低层VLA策略在50项RoboTwin 2.0任务上的平均成功率为86.20%,在4个LIBERO套件上的平均成功率为97.40%。这些结果确立了EmbodiedSkills中所用经任务适配的低层VLA策略的执行性能。在4个依赖记忆的RMBench任务上,相同的经任务适配的执行方法实现了12.5%的平均成功率。该框架提供了可训练且可检查的智能体层,用于将这些策略转化为闭环具身系统。

英文摘要

Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but long-horizon tasks require more than action prediction. An agent must coordinate perception, planning, execution, progress verification, and recovery as the physical state evolves. An action prediction or a model-generated skill decision does not, by itself, guarantee that the proposed operation is valid in the current state or that its outcome will be verified. We propose EmbodiedSkills, a unified framework that treats each skill decision as an execution proposal: the runtime checks its prerequisites before execution and verifies the outcome afterward. A shared executable-skill interface connects high-level skill selection, bounded low-level VLA execution, and post-action verification within a single agent loop. Because this interface remains fixed, low-level VLA policies can be replaced or adapted without changing the agent loop. The interface also records planning, execution, verification, and recovery events as structured trajectories, which provide supervision for individual components and can support optional online adaptation when interactive feedback is available. We instantiate EmbodiedSkills with Qwen3-VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO. Task-adapted low-level VLA policies achieve an average success rate of 86.20% across 50 RoboTwin 2.0 tasks and 97.40% across the four LIBERO suites. These results establish the execution performance of the task-adapted low-level VLA policies used in EmbodiedSkills. On four memory-dependent RMBench tasks, the same task-adapted execution approach achieves 12.5% average success. The framework provides a trainable and inspectable agent layer for turning these policies into closed-loop embodied systems.

Comments20 pages, 4 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑