发表机构
Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出仅代码策略(COAP)作为具身任务新范式,其无需 VLM/VLA,在 RoboDojo 42 项双臂任务上测试成功率达 70.24%,可支持机器人递归自我改进,还能为 VLA 等提供高效数据引擎。
AI 中文摘要
大多数机器人策略在控制循环中保留一个模型:视觉语言动作模型(VLA)将观测结果映射为动作,而智能体 harness(如策略型智能体或 harness VLA)会在运行时查询视觉语言模型(VLM)以进行决策。我们提出一种不同的视角:具身世界是一台具身图灵机,其纸带为机器人与环境状态,规则为策略。若能准确表示该状态,决策可完全用代码编写。因此我们提出仅代码策略(COAP):代码从相机图像和本体感觉中测量并跟踪机器人、环境及任务状态,并据此做出所有决策。同一代码适用于所有回合,不同任务共享一个库,控制循环中无需 VLM 或 VLA。与 VLA 和智能体 harness 相比,我们分析了 COAP 的三个优势:(i)显式状态:状态可存储在代码中;(ii)执行:代码使决策可控、灵活从故障中恢复,且在线运行快速、成本低;(iii)可扩展性:新任务可复用、继承或扩展共享库,因此能力可跨任务积累。这些优势使 COAP 成为递归自我改进(RSI)的合适媒介:编码智能体在闭环中开发库,每次变更均显式且可控。在 RoboDojo 的 42 项双臂任务上,所得库在测试时无模型的情况下达到 70.24% 的成功率。COAP 的上限取决于决策时状态表示的准确性及代码逻辑的鲁棒性。因此我们提出 COAP 作为具身任务的新范式;由于它适用于所有回合,还可作为 VLA 和智能体 harness 的高效数据引擎。
英文摘要
Most robot policies keep a model in the control loop: a VLA maps observations to actions, and an Agent Harness, such as Agent-as-Policy or Harness VLA queries a VLM for decision making at run time. We propose a different view: the embodied world is an Embodied Turing Machine, whose tape is the robot and environment state and rules are the policy. If this state can be represented accurately, the decision making can be written entirely in code. We therefore propose Code-Only-as-Policy (COAP): code measures and tracks the robot, environment, and task state from camera images and proprioception, and makes every decision from it. The same code applies across episodes, and different tasks share one library without a VLM or VLA in the loop. Compared with VLAs and Agent Harnesses, we analyze three advantages of COAP: (i) Explicit State: the state can be stored in code; (ii) Execution: code makes decision making controllable, recovers from failures flexibly, and runs fast and cheaply online; (iii) Extensibility: new tasks reuse, inherit, or extend the shared library, so capabilities can accumulate over tasks. These advantages make COAP a suitable medium for recursive self-improvement (RSI): coding agents develop the library in a closed loop, and each change is explicit and controllable. On RoboDojo's 42 bimanual tasks, the resulting library reaches a success rate of 70.24% without a model at test time. The upper bound of COAP lies in how accurately the state is represented for decision making and how robust the code logic is. We thus propose COAP as a new paradigm for embodied tasks; since it applies across episodes, it can also serve as an efficient data engine for VLAs and Agent Harnesses.
Comments31 pages, 19 figures