审计脚手架,而非检查点:智能体编码中递归自我改进的平稳性二分法
Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding
浏览论文内容
中文总结 AI 辅助
本文提出平稳性二分法,指出智能体递归自我改进受限于可达编辑集合,审计应聚焦脚手架而非检查点,并通过实验验证了收益递减与饱和现象。
中文摘要 AI 辅助
审计者若只检查系统权重是否冻结,便检查错了对象。我们的平稳性二分法指出:当智能体可达到的编辑集合保持固定时,迭代式自我修改会遭遇严格的收益递减;唯有当该集合扩张时,才能摆脱这一限制。重写脚手架(工具、验证器、分解策略)在不触碰权重的情况下扩展了智能体可达的范围,因此冻结的权重虽设定了最终上限,却无法保证过程中的平稳性。该判据还区分了通常被混为一谈的三种机制:固定类别内的搜索、提升上限本身的测试时训练,以及介于两者之间的脚手架重写。审计脚手架,而非检查点。同样的上限横向约束着其他方案。Best-of-$k$ 编排恰好实现了最佳工作者的上限:宽度带来速率,而非预算。重复咨询固定池具有可预先计算的视界,该视界仅由池本身决定;而唯一能超越它的安排——加权投票——需要真实工作者所缺乏的多样性:在30个同族工作者上,失败重叠达到最大值,多数投票在23/55(42%)的任务上失败。我们通过将细化过程解读为对草稿与目标之间残差的梯度提升(一个补丁或git差异),并测量该解读在何处失效,从而得出该判据:补丁是相互组合而非并列以供投票,且失败相互重叠。我们测量的是饱和现象。在SWE-bench上,每轮改进衰减趋近于零;在401个生产会话中,变动量呈几何级数衰减,这一形态与一个前AI人类基线共享,确立了该机制的存在却未指明其成因。这两处失效均是工程选择而非关于代码的定律,因此它们共同指明了一个值得构建的框架。
英文摘要
An auditor who checks whether a system's weights are frozen is checking the wrong thing. Our stationarity dichotomy says that iterative self-modification hits strict diminishing returns whenever the agent's reachable set of edits stays fixed, and can escape only if that set expands. Rewriting scaffolding (tools, verifiers, decomposition) expands what an agent reaches without touching a weight, so frozen weights buy an eventual ceiling but no stationarity along the way. The criterion also separates three regimes usually merged: search within a fixed class, test-time training that raises the ceiling itself, and scaffold rewriting between them. Audit the scaffold, not the checkpoint. The same ceiling binds sideways. Best-of-$k$ orchestration realizes the best worker's ceiling exactly: width buys rate, not budget. Re-consulting a fixed pool has a horizon computable in advance, decided by the pool alone, and the one arrangement that would beat it, a weighted vote, needs diversity real workers lack: on 30 same-family workers the failure overlap sits at its maximum, and a majority fails 23/55 (42%) of tasks. We obtain the criterion by reading refinement as gradient boosting on the residual error between draft and target, a patch or git diff, and then measuring where that reading breaks: patches compose instead of standing beside each other to be voted on, and failures overlap. What we measure is saturation. Per-round improvement decays toward zero on SWE-bench, and churn decays geometrically across 401 production sessions, a shape shared with a pre-AI human baseline that establishes the regime without identifying its cause. Both breaks are engineering choices rather than laws about code, so together they specify a harness worth building.
发表机构
- OnCorps
机构由 AI 辅助整理,请以论文原文为准。