arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

证明或停止:不要信任代理,信任证据——用于可验证证据门控生命周期控制的循环工程

Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control

Jek Huang, Jeffery Hsia, Jiayi Sun, Freddie Shi, Wei Huang, Ian H. White

arXiv 2607.14890首次发表:更新:

AI 中文总结

研究自主编码代理多步骤软件工作中生命周期状态缺乏证据支持的问题,提出证明或停止生命周期控制方法,经实验评估,该方法能有效控制生命周期转换,支持其作为与模型无关、主机中立的控制层。

AI 中文摘要

自主编码代理越来越多地执行多步骤软件工作,但诸如已审查、已测试、已完成和准备合并等生命周期状态除非有当前证据支持,否则仍只是声明。我们提出了证明或停止生命周期控制方法,该方法仅在新鲜的、与跟踪源状态绑定且可机械验证的证据满足相关门控时才允许生命周期转换。该方法将代理输出视为声明而非生命周期状态,并在操作上将证明定义为在既定信任模型下门控可接受的证据,而非语义程序正确性。我们通过机制测试、动力控制策略消融和操作自应用证据来评估开源实现。无人值守循环引擎在10个场景中全部通过,零误完成,本地密钥接收包拒绝了18种篡改类型,零误接受。在9240单元消融中,预注册的A4与A2'比较将可见通过/隐藏失败放大率从计算预算朴素循环下1800个注入单元中的31个降至门控循环下的2个,未放大率提高了1.6个百分点,95%置信区间为[0.8, 2.5]。接近计算量的A3与A4比较,1800个中有14个与2个,表明增益与将审查作为生命周期门控执行相关,而非仅仅添加审查者。自应用语料库包含565个故事和1007个审查结果,94.8%已解决,还有一个68行的高/关键跨供应商展示。这些结果支持证明或停止作为一个与模型无关、主机中立的控制层,用于决定生命周期可依据哪些自主代理声明采取行动。评估限于一个模型家族、24个消融任务和一个自托管语料库。

英文摘要

Autonomous coding agents increasingly execute multi-step software work, but lifecycle states such as reviewed, tested, DONE, and ready-to-merge remain claims unless supported by current evidence. We present Proof-or-Stop Lifecycle Control, a method that permits lifecycle transitions only when fresh, tracked-source-state-bound, mechanically verifiable evidence satisfies the relevant gate. The method treats agent outputs as claims rather than lifecycle state, and uses proof operationally to mean gate-admissible evidence under a stated trust model, not semantic program correctness. We evaluate an open-source implementation through mechanism tests, a powered control-policy ablation, and operated self-application evidence. The unattended-loop engine passed 10 of 10 scenarios with zero false-DONE, and local-key receipt bundles rejected 18 tamper classes with zero false accepts. In a 9,240-cell ablation, the pre-registered A4 versus A2-prime comparison reduced visible-pass/hidden-fail amplification from 31 of 1,800 injected cells under a compute-budgeted naive loop to 2 of 1,800 under the gated loop, a 1.6 percentage-point improvement in not-amplified rate with a 95 percent confidence interval of [0.8, 2.5]. A near-compute A3 versus A4 comparison, 14 of 1,800 versus 2 of 1,800, indicates that the gain is associated with enforcing review as a lifecycle gate rather than merely adding a reviewer. The self-application corpus contains 565 stories and 1,007 review findings, with 94.8 percent resolved, plus a 68-row high/critical cross-vendor exhibit. These results support Proof-or-Stop as a model-agnostic, host-neutral control layer for deciding which autonomous-agent claims a lifecycle may act on. The evaluation is limited to one model family, 24 ablation tasks, and a self-hosted corpus.

Comments48 pages, 10 figures, 29 numbered tables. Preprint v1

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑