arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自进化编码智能体:从数字程序到物理世界智能

Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence

Hongcheng Gao, Jingjing Zhou, Zelin Zheng, Shijia Ge, Jay Zhu, Yazhe Wang, Jianshu Zeng, Xuan Shangguan, Di Wu, Lingyu He, Zhiqi Jia, Sihang Wu, Xiao He

arXiv 2609.35432首次发表:更新:

AI 中文总结

提出物理编码方法,将任务状态和执行表示为代码,构建HexaAnything智能体,通过调用感知、规划和控制工具实现自进化,在RoboCasa365和物理实验中超越基线,展现从数字程序到物理世界智能的泛化能力。

AI 中文摘要

视觉-语言-动作(VLA)和世界-动作(WAM)模型直接将观测和指令映射为机器人动作。这种直接性将策略与训练绑定:轻微的布局或视角变化会导致失败,且指令泛化能力差。根本原因在于表示方式:任务要求、条件、进度和失败恢复被隐式编码在动作序列中,难以检查或修改。数字编码智能体提供了先例:大语言模型(LLM)调用工具、验证结果,并根据反馈以可执行代码的形式进行修订。显式状态、可管理的执行和可修订的程序这一相同工作模式,支撑了物理世界中的泛化和长时程执行,使物理经验得以作为可复用程序、记忆或证据返回。我们提出物理编码(Physical Coding),将任务状态和执行表示为代码。代码即世界(Code as World)记录对象、关系、约束和进度;代码即策略(Code as Policy)组织规划、验证、恢复和执行。我们构建了HexaAnything,它调用感知、规划和控制工具(包括VLA/WAM策略),并根据外部反馈进行循环内决策。验证过的轨迹成为数据和记忆,实现从工具和框架(Harness)到模型权重、架构,最终到硬件和任务设计的进化。在RoboCasa365上,HexaAnything在Composite-Unseen和总体成功率上优于XR-1 VLA,其框架训练的HexaModel在每个分割上都优于基础模型,表明代码轨迹内化了物理执行。在PhyBench和双臂AgileX机器人上,智能体自主完成物理实验和大多数桌面任务,且通常比已发表结果更快。我们观察到数据、模型和工具的自进化;未来工作目标是权重内化、架构、语言、表示和任务的自主重新设计,以及在制造业和科学领域的部署。

英文摘要

Vision-language-action (VLA) and world-action (WAM) models map observations and instructions directly to robot actions. This directness ties a policy to training: minor layout or viewpoint changes cause failure, and instructions generalize poorly. The root cause lies in representation: task requirements, conditions, progress, and failure recovery are implicitly encoded in action sequences, making them difficult to inspect or revise. Digital coding agents offer a precedent: LLMs call tools, verify results, and revise from feedback as executable code. The same working pattern of explicit state, manageable execution, and revisable procedures underlies generalization and long-horizon execution in the physical world, letting physical experience return as reusable programs, memory, or evidence. We propose Physical Coding, representing task state and execution as code. Code as World records objects, relations, constraints, and progress; Code as Policy organizes planning, verification, recovery, and execution. We build HexaAnything, which calls perception, planning, and control tools, including VLA/WAM policies, and makes in-the-loop decisions from external feedback. Verified traces become data and memory, enabling evolution from tools and Harness to model weights, architectures, and ultimately hardware and task design. On RoboCasa365, HexaAnything improves Composite-Unseen and overall success over XR-1 VLA, and its Harness-trained HexaModel beats the base on every split, indicating code traces internalize physical execution. On PhyBench and a dual-arm AgileX robot, the agent autonomously completes physics experiments and most tabletop tasks, often faster than published results. We observe data, model, and tool self-evolution; future work targets weight internalization, autonomous redesign of architectures, languages, representations, and tasks, and deployment in manufacturing and science.

CommentsTechnical report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑