在启发式自我改进智能体中自我编写的验证不可靠
Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents
AI总结:
研究启发式自我改进智能体中自我编写验证不可靠的问题,引入密封外部接受循环(SEAL)方法,通过实验发现该问题在启发式学习中常现,自我验证失败按能力分层,SEAL性能优于无保护基线。
AI中文摘要:
自我改进智能体通过反复重写程序策略、控制器或启发式规则来积累能力,通常依靠自我编写的测试或指标来决定是否接受后续编辑。由于智能体同时控制优化对象及其验证器,会出现自我评分高但实际部署性能下降的情况。本文通过验证器-部署差距来研究该问题,询问自我编写的验证在迭代策略和测试重写中如何失败等。为此引入密封外部接受循环(SEAL),实验表明该问题在启发式学习设置中常出现,自我编写验证的失败按能力分层,SEAL在六个模型和三个随机种子上优于无保护基线。
英文摘要:
Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-authored tests or metrics to decide whether to accept subsequent edits. The agent controls both the optimized object and its verifier. As a result, self-assigned scores can remain near perfect while real deployment performance degrades or stays low. We study this problem through the verifier--deployment gap. This gap refers to the discrepancy between an agent's self-authored verification signal and a sealed deployment evaluation that the agent cannot observe or access. We ask how self-authored verification fails under iterative policy-and-test rewriting, how the failure changes with capability, and how little exogenous trust is sufficient to prevent real regressions from being deployed. To address this problem, we introduce a Sealed Exogenous Acceptance Loop (SEAL). SEAL retains self-authored tests but compares each candidate with the incumbent through a fixed harness-side audit. The agent cannot author or inspect the audit, receives only accept/reject, and the whole incumbent state is retained after a clear regression. Our experiments show that this problem often appears in heuristic learning settings. These settings require trial-and-error discovery of the target objective. We further find that failures of self-written verification are stratified by capability. Weaker agents tend to damage previously acquired strategies behind easy self-tests. Stronger agents are more stable, but they still mismeasure the deployment distribution. Standard self-written constraints do not reliably close this gap. In contrast, SEAL outperforms unprotected baselines across six models and three random seeds. Reliable self-improvement need not abandon self-verification, but it requires at least one deployment-acceptance signal outside the agent's control.