arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10157cs.AI

SBCO:面向规划智能体的自监督、验证器驱动的工具链优化器

SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning Agents

Vivek Kulkarni, Sudipta Paul, Aounon Kumar, Nicholas Tzou, Srinivas Chappidi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对规划任务提出SBCO,一种自监督的验证器驱动工具链优化器,其性能优于或相当的定制基线,且计算成本仅为其1/4至1/5.5。

中文摘要 AI 辅助

自我改进智能体旨在通过让AI系统随时间演化并提升性能,减少人类在AI系统背后的工程工作量。近期,Darwin Gödel Machine和Huxley Gödel Machine等方法被提出,这些方法通过自引用实现开放式、递归的自我改进,其中编码智能体可编辑自身代码。这类自引用自我改进方法要求执行任务所需的能力与自我修改所需的能力一致或高度对齐,编码任务满足这一条件。对于不满足所需对齐条件的领域或任务,无法采用自引用自我改进。这种情况下,可通过去除自引用方面或引入元智能体的显式自我修改来将上述算法适配到其他任务,但这两种方式计算成本高昂,依赖于对大量候选智能体的种群或自我修改搜索。对于具有显式约束的规划任务,我们提出一种成本低得多的替代方案。我们引入SBCO(自监督块坐标优化器,Self-supervised Block Coordinate Optimizer),它是一种验证器驱动的工具链优化器,属于与Gödel机方法同属闭环、基于经验改进的系列,但采用自监督而非自引用机制。给定一个智能体工具链,SBCO通过近似块坐标上升学习分解后的验证器库和工具链策略,利用自身的分级反馈改进智能体的输出——使用固定的元智能体,无需人工标签。在两个领域中,SBCO的性能与定制的自我修改基线相当或更优,同时计算预算减少4至5.5倍。

英文摘要

Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over time. Recently, methods like the Darwin Gödel Machine and the Huxley Gödel Machine have been proposed which enable open-ended, recursive self-improvement through self-reference where a coding agent edits its own code. Such self-referential self-improvement methods require that the competence required to perform the task coincides or aligns well with the competence required for self-modification which is the case for coding tasks. For domains or tasks, which do not satisfy the alignment needed, self-referential self-improvement is not available. In such cases, it is possible to adapt the above algorithms to other tasks by removing the self-referential aspect or introducing explicit self-modification of a meta-agent -- both computationally expensive, relying on population or self-modification search over many candidate agents. For planning tasks with explicit constraints, we propose a far cheaper alternative. We introduce SBCO (Self-supervised Block Coordinate Optimizer), a verifier-grounded harness optimizer in the same closed-loop, improve-from-experience family as the Gödel-machine methods, but self-supervised rather than self-referential. Given an agentic harness, SBCO learns a decomposed bank of verifiers and a harness policy via approximate block coordinate ascent, improving the agent's outputs from its own graded feedback---with a fixed meta-agent and no human labels. Across two domains SBCO matches or exceeds a customized self-modifying baseline while using 4-5.5 times less compute budget.

↑