AI 中文总结
该研究将编码智能体的软件循环架构迁移至机器人策略改进领域,提出AgenticRobotics控制平面,解决机器人工具易失效问题,实现错误提升控制等性能提升。
AI 中文摘要
Claude Code和Codex等编码智能体可关闭软件循环:主智能体管理循环,子智能体负责分析与执行,工具完成具体工作。我们将此架构迁移至机器人策略改进领域,设计中存在一个关键差异:机器人工具——包括训练好的策略、训练流水线、数据收集模块——会频繁失效,因此必须对工具质量进行评估,在每次调用时记录相关信息,并在其背后的工件发生变化时将其标记为失效。AgenticRobotics是一个与后端无关的控制平面,其中大型语言模型(LLM)控制器通过持久的训练-评估-改进事务驱动一次性工作者:包含不可变的目标、由控制器掌控的测量机制、基于提交密钥的崩溃恢复、按证据分级的技能库,以及带有标准化、可记录调用接口的工具注册表。标题是一个可操作的主张,而非选择主张:操作人员可以离开,因为能力提升受证据管控,状态可恢复,且能力质量源自记录——而非因为该循环比人类能选出更好的检查点;在我们测量的一条谱系上,它并未做到这一点。这些管控措施可显著提升错误提升的控制能力(每轮运行的错误率经强化后为0.001,而交付版本为0.005至0.021),支持在可选停止条件下随时做出有效决策,在终止注入场景下无效果丢失或重复,且由签名验证器捕获了全部6类工件篡改情况。
英文摘要
Coding agents such as Claude Code and Codex close the software loop: a main agent manages the loop, subagents analyze and execute, tools do the work. We port this architecture to robot-policy improvement, where one difference dominates the design: robotic tools---trained policies, training pipelines, data collection---fail routinely, so a tool's quality must be measured, recorded at every call, and expired when the artifact behind it changes. AgenticRobotics is a backend-independent control plane in which an LLM controller drives disposable workers through durable train--evaluate--improve transactions: an immutable objective, controller-owned measurement, commit-keyed crash recovery, an evidence-graded skill library, and a tool registry with a standardized, recorded call surface. The title is an operational claim, not a selection claim: the operator can leave because promotion is evidence-gated, state is recoverable, and capability quality is derived from records---not because the loop picks better checkpoints than a human; on the one lineage we measured, it does not. The gates measurably buy false-promotion control (0.001 per run hardened versus 0.005--0.021 shipped), anytime-valid decisions under optional stopping, zero lost or duplicate effects under kill injection, and six of six artifact-tampering classes caught by a signed verifier.
CommentsAgentic Robotics Preview Version