arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

让代码即策略再次伟大:前沿智能体编写、调用并进化机器人工具

Make Code as Policy Great Again: Frontier Agents Write, Call, and Evolve Robot Tools

Shijia Ge, Alex Zhou, Jianshu Zeng, Yexing Wan, Di Wu, Zelin Zheng, Yazhe Wang, Zhiqi Jia, Xuan Shangguan, Jay Zhu, Yijun Liu, Lingyu He, Sihang Wu, Xiao He, Hongcheng Gao

arXiv 2609.39018首次发表:更新:

发表机构

Hexafuture Inc.; Peking University; Beijing Institute of Technology; Tsinghua University(Hexafuture 公司; 北京大学; 北京理工大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

URAI通过编程智能体构建可执行工具、执行智能体调用工具,将前沿模型的控制决策与代码执行分离,在RoboDojo和真实双臂任务中显著提升成功率与效率。

AI 中文摘要

前沿模型能够控制机器人,但通过推理每一次伸展、抓取和后退会使操作变得缓慢且耗费大量令牌。我们重新审视代码即策略,采用不同的分工方式:模型构建可执行工具,代码处理多阶段运动,模型决定下一步做什么。我们引入URAI(通用机器人-智能体接口),它将构建机器人工具的编程智能体与在反馈循环中使用这些工具的执行智能体相结合。编程智能体根据任务意图编写可重用且特定于任务的工具,并通过执行反馈和人类指导对其进行改进。执行智能体根据当前观测选择并参数化这些工具;每次调用在将控制权返回给智能体之前,在本地运行完整的运动。与将后续决策委托给生成的程序不同,这种设计在工具执行之间保留了模型级别的决策能力。经过验证的工具修订会在多个回合中持续存在,而无需更新基础模型权重,并且共享的GUI和API使相同的工具可供人类和智能体使用。在五个RoboDojo任务和四个冻结的执行智能体上,相对于直接指尖控制,URAI将总体成功率从18.0%提升至53.0%,其中在Swap Blocks上提升最大;使用相同的工具,预先编写的程序仅达到24%,而两个在每个调用后决策的智能体达到56%。四个智能体中有三个还以1.3-1.5倍的速度完成回合,执行智能体输出令牌减少1.5-1.7倍;DeepSeek-V4-Flash的成本几乎不变。我们进一步在七个真实世界的AgileX双臂任务上评估URAI,涵盖物体操作、布料折叠和人机交互的井字棋。URAI连接了前沿智能体的编码和决策能力,围绕可重用工具组织机器人控制,智能体既可以调用这些工具,也可以对其进行修订。

英文摘要

Frontier models can control robots, but reasoning through every reach, grasp, and retreat makes manipulation slow and token-intensive. We revisit code as policy with a different division of labor: models build executable tools, code handles multi-phase motions, and models decide what to do next. We introduce URAI (Universal Robot-Agent Interface), which couples a programming agent that constructs robot tools with an execution agent that uses them in a feedback loop. The programming agent writes reusable and task-specific tools from task intent and refines them through execution feedback and human guidance. The execution agent selects and parameterizes these tools from current observations; each call runs a complete motion locally before returning control to the agent. Unlike delegating subsequent decisions to a generated program, this design retains model-level decision-making between tool executions. Validated tool revisions persist across episodes without updating foundation-model weights, and a shared GUI and API make the same tools available to humans and agents. Across five RoboDojo tasks and four frozen execution agents, URAI raises aggregate success from 18.0% to 53.0% relative to direct fingertip control, with the largest gain on Swap Blocks; with the same tools, a program written in advance reaches only 24% against 56% for two agents deciding after each call. Three of the four agents also finish episodes 1.3-1.5 times faster with 1.5-1.7 times fewer execution-agent output tokens; DeepSeek-V4-Flash's cost barely changes. We further evaluate URAI on seven real-world AgileX dual-arm tasks, spanning object manipulation, cloth folding, and human-interactive tic-tac-toe. URAI connects the coding and decision-making capabilities of frontier agents, organizing robot control around reusable tools that agents can both invoke and revise.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑