arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大模型提出方案,执行体处理:一种自验证智能体工具,可将长程智能体中的承诺漂移与绑定漂移分离

The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents

Mohsen Arjmandi

arXiv 2608.04066首次发表:更新:

AI 中文总结

本研究提出一种自验证智能体工具,可分离长程智能体的承诺漂移与绑定漂移,通过实验揭示消融承诺机制会提升目标放弃率,且贡献了可量化漂移分解的智能体验证方法论。

AI 中文摘要

当长程智能体的自身状态和自我报告完全不可信时,如何对其进行验证?我们提出一种智能体工具,其验证是结构性的而非事后的。确定性执行体拥有全部信念;语言模型仅可提交带类型的提案,且只有当行动前预先注册的预测与代码观测结果匹配时,主张才会被认可。该工具具备两项特性,使其不仅是智能体的验证器,更是自身科学的验证器:每次运行若出现按组织划分的写入错误、渲染尺寸或盐值金丝雀回显下限被突破,便会自动失效(前8次架构运行中有4次失效,每次都定位到了真实缺陷);不可见渲染的影子参考会编译出完整系统在每个消融单元中本应执行的计划,因此即便被测机制已被移除,仍可定义漂移指标。利用该工具,我们报告了关于长程智能体普遍存在的故障的清晰单变量结果:消融承诺机制会将目标放弃率从0.00翻转为1.00,而绑定误差保持0.00不变(每个单元3个种子,每次运行最多394个参考节拍,所有运行均为有效门控)。相比之下,当消融绑定通道的修复机制时,其不会作为每个节拍的漂移重新出现——由于绑定由代码掌控,故障类别被结构性吸收,唯一残留会在上游一层表现为假设形成的崩溃。我们完全公开任务效能为零(在ARC-AGI-3的52次门控运行中完成率为0),这一情况已预先注册为结构性失败。本研究的贡献是一种用于智能体开发的验证方法论,以及该方法论可量化的漂移分解。

英文摘要

How do you verify a long-horizon agent when its own state and self-reports are exactly what you cannot trust? We present an agent instrument built so that verification is structural rather than post-hoc. A deterministic Executive owns all belief; a language model may only file typed proposals, and a claim is admitted only when a prediction pre-registered before acting is matched against observation by code. Two properties make the instrument a verifier of its own science, not just of the agent: every run invalidates itself when per-organ write-error, render-size, or salted-canary-echo floors are breached (four of the first eight architecture runs were invalidated, each localizing a real defect); and a render-invisible shadow reference compiles the plan the full system would have committed in every ablation cell, so drift metrics are defined even where the mechanism under test has been removed. Using this instrument we report a clean, single-variable result on a failure every long-horizon agent suffers: ablating the commitment mechanism flips goal-abandonment from 0.00 to 1.00 while binding error stays flat at 0.00 (three seeds per cell, up to 394 reference beats per run, every run gated valid). The binding channel, by contrast, does not reappear as per-beat drift when its repair is ablated -- because binding is code-owned, the failure class is structurally absorbed, its only residue appearing one layer upstream as a collapse in hypothesis formation. We report these under full disclosure that task efficacy is null (zero level completions across 52 gated runs on ARC-AGI-3), pre-registered as a structural defeater. The contribution is a verification methodology for agent development and the drift decomposition it makes measurable.

Comments7 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑