发表机构
Novaxbot(Novaxbot)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明视觉界面机器人智能体的工具链可由优化器智能体自动进化,其效果取决于每轮回滚规模(增大批次提升信噪比)和修订选择机制(冠军-挑战者机制防止性能退化),可将留出成功率从51%提升至67%。
AI 中文摘要
当现成的编码智能体被直接用作机器人策略时,通过截图观察基于浏览器的3D界面,并借助少量工具通过虚拟目标夹爪进行动作,该智能体的工具链(即其提示词、工具和控制规则)在很大程度上决定了成败,而迄今为止这些工具链都是手工编写的。我们证明,这种工具链可以由另一个编码智能体(即优化器智能体)自动改进,并报告了关于其有效性的两项发现。首先,优化器智能体每轮观察到的回滚次数决定了进化后的工具链是否可信、是否具有泛化能力以及是否能够稳定改进。单次回滚是一个带有噪声的二元结果,因此当每轮回滚次数较少时,修订可能仅凭运气被提升;增大批次规模可以提高每次提升决策的信噪比。在保持轮数固定的情况下,将训练集从5次回滚增加到100次回滚,留出成功率从47%上升到67%,而较小的训练集则会出现过拟合,在训练任务上达到70%,但留出成功率仅为54%。其次,优化器智能体不能被赋予完全的自由。如果其提出的每次修订都被无条件接受,那么由于不明智的编辑不断累积,性能会在十轮内持续下降;而加入最基本的保障措施——冠军-挑战者选择机制,即仅当修订在相同的固定评估集上严格优于当前版本时才予以提升——则会将同一循环转变为在30轮内将留出成功率从51%提升到67%的过程。因此,针对视觉界面机器人智能体的自动工具链进化是可行的,但其收益取决于每次决策背后的回滚规模以及优化器智能体的修订如何被选择。
英文摘要
When an off-the-shelf coding agent is used directly as a robot policy, observing a browser-based 3D interface through screenshots and acting by posing a virtual target gripper through a few tools, the agent's harness, its prompts, tools, and control rules, largely determines success, and until now it has been written by hand. We show that this harness can be improved automatically by another coding agent, the optimizer agent, and report two findings about what makes it work. First, the number of rollouts the optimizer agent sees per round governs whether the evolved harness is trustworthy, generalizes, and improves steadily. A single rollout is a noisy binary outcome, so with few rollouts per round a revision can be promoted on luck; enlarging the batch raises the signal-to-noise ratio of every promotion decision. Holding rounds fixed and growing the training set from 5 to 100 rollouts, held-out success rises from 47% to 67%, while small training sets overfit, reaching 70% on training tasks but only 54% held-out. Second, the optimizer agent must not be given free rein. With every revision it proposes accepted unconditionally, performance drifts downward within ten rounds as ill-judged edits accumulate; adding the most basic safeguard, Champion-Challenger selection that promotes a revision only if it strictly beats the incumbent on the same fixed evaluation set, turns the same loop into one that raises held-out success from 51% to 67% over 30 rounds. Automatic harness evolution for visual-interface robot agents is thus feasible, but its gains hinge on the rollout scale behind each decision and on how the optimizer agent's revisions are selected.
Comments12 pages, 4 figures