arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体循环何时会将停滞误认为进步?长期运行的自主语言模型智能体循环中的自我评估偏差与外部基础验证

When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops

Hyundoo Park, Byungho Choi

arXiv 2607.25152首次发表:更新:

AI 中文总结

研究长期运行的自主智能体循环中自我评估偏差导致的“进步幻觉”问题,通过构建测试平台,控制评估者信息通道类型,发现自我报告无意义,带内评判者也有偏差,指出开放式目标需带外评估,强调评估者依据对结果的重要性。

AI 中文摘要

长期运行的自主智能体在没有人工干预的情况下进行规划、行动并判断自身是否完成任务。当智能体对自己的工作进行评分时,自我评估偏差就会出现:看似合理的变化被视为进步,而实际结果却停滞不前或出现倒退。我们将这种失败模式称为“进步幻觉”,并通过控制测量表明,这是一个评估者基于何种依据的问题。我们构建了一个测试平台,固定智能体及其工具表面,仅操纵控制循环的评估者的信息通道类型。通过容器和网络隔离实现了原则上不可伪造的世界状态预言机,并在每次运行时进行验证。在54个周期中,一个前沿智能体每次都声称有所改进,但56%的测量变化量为零或更低。因此,自我报告毫无意义,自我裁决门退化为全部接受,使其达到的最佳部署状态下降了19%。即使是最强的带内评判者,阅读完整的工件文本、更改差异及其自己的裁决历史,也接受了44%实际是现实世界倒退的周期,并拒绝了38%的实际改进;预先注册的对抗性假设,即强大的评判者可以缩小差距,被拒绝。在一个成功规范可从工件本身验证的边界任务中,同一位评判者的幻觉消失为零,差距在注册阈值内缩小,表明差距取决于成功信号所在的位置。一个仅返回接受裁决的仅符号变体使现实世界输出与完整反馈相似(110.0对113.0),表明好处在于门的基础而非反馈内容。对于成功信号存在于记录之外的开放式目标,扩大评判者规模是不够的;具有现实世界访问权限的带外评估是一项结构性要求。

英文摘要

Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation bias takes hold: plausible changes are accepted as progress while real-world outcomes stagnate or regress. We name this failure mode the progress mirage and show, with controlled measurement, that it is a question of what the evaluator is grounded in. We built a testbed that holds the agent and its tool surface fixed and manipulates only the information-channel type of the evaluator that gates the loop. A world-state oracle, unfakeable in principle, is enforced by container and network isolation and verified at every run. Across 54 cycles a frontier agent claimed improvement every time, yet 56 percent had a measured delta of zero or below. Self-report was thus uninformative, and the self-verdict gate degenerated into accept-all, eroding the best deployed state it had reached by 19 percent. Even the strongest in-band judge, reading the full artifact text, the change diff, and its own verdict history, accepted cycles of which 44 percent were real-world regressions and rejected 38 percent of real improvements; the preregistered adversarial hypothesis that a strong judge closes the gap was rejected. On a boundary task whose success specification is verifiable from the artifact itself, the same judge's mirage vanished to zero and the gap collapsed within the registered threshold, showing that the gap depends on where the success signal resides. A sign-only variant returning only the acceptance verdict kept real-world output similar to full feedback (110.0 versus 113.0), locating the benefit in the gate's grounding rather than in feedback content. For open-ended objectives whose success signal lives outside the transcript, scaling up the judge is not enough; out-of-band evaluation with real-world access is a structural requirement.

Comments23 pages. Preregistered pilot measurement study. Also deposited on Zenodo (concept DOI 10.5281/zenodo.21594735)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑