发表机构
Google; OMNI3ai; Glinr Studios(谷歌; OMNI3ai; Glinr Studios)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Assay通过将智能体主张绑定到代码依赖锥的Merkle哈希,实现主张随代码变化而精确失效,并引入风险成比例的证据义务与合并门,以低成本确保AI辅助软件交付的可问责性。
AI 中文摘要
AI编码智能体以两种耦合的方式失败。它们花费大部分上下文窗口重新发现事物所在位置,并且在工作变得困难时在没有证据的情况下断言成功。仓库索引通过廉价上下文解决前者,而具有对抗性审查的编排框架通过问责制解决后者。两者都描述了同一对象,即代码库的结构,在两个时间尺度上:代码当前的真实情况,以及在哪个修订版、由谁验证为真实的情况。Assay使这一观察变得可操作。智能体做出的每一个主张(测试通过、无秘密、行为保持)都绑定到其所覆盖代码的依赖锥的Merkle哈希,因此当该代码或其依赖的任何内容发生变化时,该主张恰好变得过时。我们证明了该绑定是健全且最小的,并且变更的爆炸半径恰好是它使其失效的主张集合。在图上,我们放置了风险成比例的证据义务、具有职责分离的有界审查协议,以及一个不咨询任何模型的合并门:覆盖率、新鲜度、签名、退出代码、合理性、证据单调性(“不要删除失败的测试”的机械形式)和审查状态。Assay是一个无依赖的Python工具,带有MCP服务器。在五个公共仓库上,600令牌的简报比探索代理便宜14倍到114倍,热重建比冷重建快多达5倍,锥绑定重新验证了7.9%到81.9%的主张,而仓库绑定重新验证所有主张,而每模块绑定则遗漏了23%到68%的必要失效,该门阻止了9个脚本化对抗行为中的9个,同时允许诚实的行为。本文中的每个数字都由发布的脚本生成。
英文摘要
AI coding agents fail in two coupled ways. They spend most of their context window rediscovering where things live, and they assert success without evidence when the work gets hard. Repository indexes address the first with cheap context, and orchestration frameworks with adversarial review address the second with accountability. Both describe the same object, the structure of the codebase, at two timescales: what is true of the code now, and what was verified to be true, at which revision, by whom. Assay makes that observation operational. Every claim an agent makes (tests pass, no secrets, behavior preserved) is bound to the Merkle hash of the dependency cone of the code it covers, so the claim is stale exactly when that code or anything it depends on changes. We show the binding is sound and minimal, and that the blast radius of a change is precisely the set of claims it invalidates. On the graph we place a risk-proportional evidence obligation, a bounded review protocol with separation of duties, and a merge gate that consults no model: coverage, freshness, signatures, exit codes, plausibility, evidence monotonicity (the mechanical form of "do not delete the failing test"), and review status. Assay is a dependency-free Python tool with an MCP server. On five public repositories a 600-token brief costs 14x to 114x less than an exploration proxy, warm rebuilds are up to 5x faster than cold ones, cone binding re-verifies 7.9% to 81.9% of claims where repository binding re-verifies all of them while per-module binding misses 23% to 68% of required invalidations, and the gate blocks 9 of 9 scripted adversarial behaviours while admitting the honest ones. Every number in this paper is generated by the released scripts.
Comments8 pages, 4 figures, 5 tables. Code, experiments, and paper source: https://github.com/OmShiv/assay-research