arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25130cs.CLcs.AIcs.LG

影响并非失效:应询问声明而非差异

Impact Is Not Invalidation: Ask About the Claim, Not the Diff

  • Thomson Reuters(汤森路透)

机构由 AI 辅助整理,请以论文原文为准。

Atul Anand

AI总结:

针对编码智能体记忆系统,本文发现询问差异是否保持行为不如询问具体声明是否成立,后者精确度大幅提升,并构建了基于执行验证的翻转声明数据集。

AI中文摘要:

编码智能体的记忆系统必须在仓库发生变化时,决定其存储的哪些声明已变为虚假。内容锚定法在声明来源的工件发生改变时即判定该声明失效,这会导致频繁触发。语义等价分类则询问差异是否保持行为不变,这是一个关于差异而非任何存储声明的问题。我们证明第二种信号因与模型能力无关的原因而失败:当被问及一次提交是否保持行为时,跨越40倍价格区间的五个模型在59-72%的真实提交上触发,且相对于0.25的基线比率,精确度仅为0.291至0.329。而当被问及某一特定声明是否仍然成立时,相同模型在相同差异上达到0.705至0.974的精确度。一个对照组将行为保持判定器的输入加上声明文本,仅改变问题,精确度变化为0.010和0.016;而改变问题则使精确度变化0.49和0.65。我们还与pytest-testmon(一个已部署的回归测试选择器,具有基于覆盖率的依赖数据)进行比较:它在0.415的精确度下达到0.868的召回率,因此对变更可能触及内容的近乎完全了解并不能识别其使哪些内容失效。真实标签是执行而非注释:声明是在提交t时通过的测试函数,如果相同断言文本在t+1时失败,则该声明已翻转。构建此数据集需要我们在先前工作中未发现的观察。在CI门控的主线上,使既有测试失败的提交无法合并,因此朴素构造的设计中正类为空。我们报告了从23个Python库中挖掘出的10,369条声明,其中184条经执行验证的翻转,按仓库划分的保留集、知识截止后的划分、打乱差异的零假设、释义对照组,以及跨越17个仓库的留一仓库分析。

英文摘要:

Memory systems for coding agents must decide, when a repository changes, which of their stored claims have become false. Content anchoring invalidates a claim whenever the artifact it came from changes, which fires constantly. Semantic-equivalence classification asks whether a diff preserves behavior, a question about the diff rather than about any stored claim. We show the second signal fails for a reason unrelated to model capability: asked whether a commit preserves behavior, five models spanning a 40x price range fire on 59-72% of real commits and reach precisions of only 0.291 to 0.329 against a 0.25 base rate. Asked instead whether one specific claim still holds, the same models on the same diffs reach 0.705 to 0.974. A control that hands the behavior-preservation judge the claim text, changing only the question, moves precision by 0.010 and 0.016; changing the question moves it by 0.49 and 0.65. We also compare against pytest-testmon, a deployed regression-test selector with coverage-derived dependency data: it reaches 0.868 recall at 0.415 precision, so near-complete knowledge of what a change can reach does not identify what it falsifies. Ground truth is execution, not annotation: a claim is a test function passing at commit t, and it has flipped if that same assertion text fails at t+1. Building this required an observation we did not find in prior work. On a CI-gated mainline a commit that leaves a pre-existing test failing cannot merge, so the naive construction has an empty positive class by design. We report 10,369 claims with 184 execution-verified flips mined from 23 Python libraries, splits held out by repository, a post-knowledge-cutoff split, a shuffled-diff null, a paraphrase control, and a leave-one-repository-out analysis over 17 repositories.

补充信息

↑