发表机构
Max Planck Institute for Software Systems; EPFL; Apple; Aarhus University(马克斯·普朗克软件系统研究所; 洛桑联邦理工学院; 苹果公司; 奥胡斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对仓库规模编码任务,建模一致性债务,通过7个模型和5个框架实验发现事实可用性决定编码结果,提出测试框架应核对编辑依赖事实可用性的建议。
AI 中文摘要
仓库规模的编码要求智能体在有限的上下文窗口内保持测试、导入、配置和迁移规则的一致性。我们将此建模为耦合事实图的重构:在每次编辑时,所需事实来自近期上下文或参数记忆,而未被两者覆盖的事实则构成一致性债务。我们在7个模型和5个测试框架中分别提供和 withhold 各通道,并注入故障。正如预期,当两个通道均为空时,没有模型能在未见过的API上完成任务,而将事实放入提示中可恢复成功。当重命名操作破坏了模型对真实库的记忆时,所有7个模型都会在同一位置失败,通过和遗漏相同的测试。可用性决定结果,而距离不决定: withhold 一个事实会恰好损失它所支持的工作量,且提供的事实在离编辑较远的位置与紧邻编辑的位置效果相同。测试框架为此付出不等的代价:所有测试都通过的配置在消耗的token数量上差异超过十倍,因为它们以不同速率重建相同内容,且当事实被 withhold 时,花费更多也无济于事。缺失的事实会产生错误的工作而非缺失的工作:被要求行动的智能体会执行操作,编造文件或猜测值,因此基于读取构建的工具会寻找一个已被填补的漏洞。智能体表示被阻止的频率是模型的属性,从每次试验都表示到无一次表示不等。可用性并非决定每次编辑:当标准与代码不一致时,智能体会遵循标准,即使标准规定的代码更差,因此过时的约定文件比没有文件代价更高。由于参数记忆替代了读取,在SWE-bench上,模型可能了解仓库,读取不再能预测成功。测试框架应在智能体写入时保持编辑依赖的事实可用,并将此可用性与智能体生成的内容而非其读取的内容进行核对。
英文摘要
Repository-scale coding requires an agent to keep tests, imports, configuration, and migration rules consistent within a bounded context window. We model this as reconstructing a coupled-fact graph: at each edit, a required fact comes from recent context or parametric memory, and the facts covered by neither form coherence debt. We supply and withhold each channel and inject faults across seven models and five harnesses. As expected, no model completes a task on an unseen API with both channels empty, and putting the facts in the prompt restores success. When a rename defeats what models memorized about a real library, all seven fail in the same place, passing and missing the same tests. Availability decides the outcome and distance does not: withholding a fact costs exactly the work it supports, and a supplied fact works as well far from the edit as next to it. Harnesses pay unequal prices for it: configurations that all pass every test differ more than tenfold in tokens consumed because they rebuild the same content at different rates, and spending more recovers nothing when facts are withheld. A missing fact produces wrong work rather than absent work: an agent asked to act acts, fabricating the file or guessing the value, so instruments built on reads look for a hole already filled. How often it says it is blocked instead is a property of the model, from every trial to none. Availability does not settle every edit: where standard and code disagree, agents follow the standard even when it prescribes the worse code, so a stale convention file costs more than no file. Because parametric memory substitutes for reading, on SWE-bench, where models likely know the repositories, reads no longer predict success. Harnesses should keep the facts an edit depends on available when the agent writes, and check that availability against what the agent produces rather than what it reads.