arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04278cs.SEcs.AI

EA-Graph:面向上游漂移场景下编码智能体的工件锚定验证记忆

EA-Graph: Artifact-Anchored Verification Memory for Coding Agents under Upstream Drift

  • National Yang Ming Chiao Tung University(国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Hwai-Jung Hsu, Cheng-Jan Chi, Hanna Everett

AI总结:

该研究提出EA-Graph这一工件锚定验证记忆,经实验证实其在编码智能体的可证性判断任务中,能提升较小模型的表现,可缩小模型间能力差距。

AI中文摘要:

编码智能体日益在多会话中工作,但文本笔记可保留结论却不记录支撑该结论的程序状态。上游变更后,仓库可能仍能构建,即便此前的验证主张已无效。EA-Graph是一种用于验证主张的工件锚定记忆,它以子路径粒度表示工件,将别名解析为叶子定义,将每个主张锚定到建立该主张所用的内容,并将证据强度与新鲜度分离。当替换内容不可用时,该主张变为不可证,而非猜测。EA-Graph在生成式仓库上进行评估,这些仓库的行为与工件的对应关系是通过构造已知的。任务是在值漂移、逻辑漂移和故意被扣留的上游内容后,将先前的主张分类为未受影响、受影响或不可证。分析覆盖了7个干净环境中的42个会话、14个模型-环境实例、3种记忆条件和2个模型层级。在Haiku轮次中,工件锚定记忆在全部7个环境中均优于文本笔记和无持久记忆;每一次精确配对的Wilcoxon检验均得出p=0.0156。在Sonnet轮次中,锚定条件表现完美,但频繁的控制上限使得预注册的对比不显著。没有会话捏造被扣留的内容。这些结果支持一个有限结论:在该测试平台中,工件锚定记忆提升了较小模型的可证性判断。一项探索性对比进一步表明,结构化主张记忆可通过外化会话内的重新推导缩小能力差距,但并未确立跨模型等价性。该研究未对效率或修复质量提出主张。

英文摘要:

Coding agents increasingly work across sessions, but prose notes can preserve a conclusion without the program state that supported it. After an upstream change, a repository may still build even though earlier verification claims are no longer valid. EA-Graph is an artifact-anchored memory for verification claims. It represents artifacts at sub-path granularity, resolves aliases to leaf definitions, anchors each claim to the content used to establish it, and keeps evidence strength separate from freshness. When replacement content is unavailable, the claim becomes unprovable rather than guessed. EA-Graph is evaluated on generated repositories whose behavior-to-artifact ground truth is known by construction. The task is to classify prior claims as unaffected, affected, or unprovable after value drift, logic drift, and deliberately withheld upstream content. The analysis covers 42 sessions across seven clean worlds, 14 model-world instances, three memory conditions, and two model tiers. In the Haiku round, artifact-anchored memory outscored prose notes and no persistent memory in all seven worlds; each exact paired Wilcoxon comparison yielded p = 0.0156. In the Sonnet round, the anchored condition was perfect, but frequent control ceilings left the preregistered contrasts non-significant. No session fabricated withheld content. These results support a bounded claim: artifact-anchored memory improved the smaller model's provability judgments in this testbed. An exploratory comparison further suggests that structured claim memory may narrow a capability gap by externalizing in-session re-derivation, but it does not establish cross- model equivalence. The study makes no claim about efficiency or repair quality.

补充信息

↑