arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

真实软件历史中的时间有效性:基于GitHub修复消除代码助手记忆中的过时事实错误

Temporal Validity on Real Software Histories: Eliminating Stale-Fact Errors in Code-Assistant Memory over GitHub Fixes

Neeraj Yadav

arXiv 2608.20685首次发表:更新:

发表机构

Called It Inc.(Called It 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对代码助手RAG存在的无法区分新旧代码事实的问题,提出MemStrata确定性取代记忆,在真实GitHub修复数据集上使过时事实错误率降至约0,准确率达0.91,延迟仅约2.1秒。

AI 中文摘要

检索增强生成(RAG)没有时间模型:当编码会话中的事实发生变化时——函数被重命名、端点被移动、依赖项被更新——RAG会以几乎相同的相似度检索到旧值和新值,无法区分哪个是当前值,因此会提供已被取代的值。论文1在合成单值基准测试中表明,确定性的(主体、关系、客体)取代记忆可消除这种故障。在此,我们在真实软件历史上对其进行端到端验证。从707个真实GitHub问题(SWE-bench Lite + Verified)中,我们提取了130个干净的原子状态转换,即从修复前形式到修复后形式更改一个可识别值的修复,且每个转换都是无标记的(过时语句和当前语句仅在值上不同)。在该数据集上,MemStrata的答案准确率达到0.91,而RAG的准确率为0.57-0.59;此外,当被强制回答时,RAG有36-38%的时间会提供已被取代的值(LLM重排序器对此无帮助),而MemStrata将这一比例降至约0,且延迟与RAG检索延迟相近(约2.1秒,而重排序器约为18秒)。我们明确说明适用范围:仅约18%的真实修复是干净的原子转换;论文2针对该类问题分离出记忆机制,其余修复的提取覆盖是正交问题,我们将其留待后续工作解决。研究期间还出现并修复了一个真实产品错误(不区分大小写的值比较/标点符号),且确定性取代在干净代码变异上的准确性(护城河属性)得以保留和验证。

英文摘要

Retrieval-augmented generation (RAG) has no model of time: when a fact changes across a coding session - a function is renamed, an endpoint moves, a dependency is bumped - RAG retrieves both the old and new value with near-identical similarity and cannot tell which is current, so it serves the superseded value. Paper 1 showed, on synthetic single-value benchmarks, that a deterministic (subject, relation, object) supersession memory eliminates this failure. Here we validate it end-to-end on real software history. From 707 real GitHub issues (SWE-bench Lite + Verified) we extract 130 clean atomic state transitions, a fix that changes one identifiable value from a pre-fix to a post-fix form, and render each marker-free (the stale and current statements differ only in the value). On this set, MemStrata reaches 0.91 answer accuracy versus RAG's 0.57-0.59; and, the structural result, when forced to answer RAG serves the superseded value 36-38% of the time (an LLM reranker does not help) while MemStrata drives this to ~0, at RAG retrieval latency (~2.1 s vs ~18 s for the reranker). We are explicit about scope: only ~18% of real fixes are clean atomic transitions; Paper 2 isolates the memory mechanism on that class, and extraction coverage of the remaining fixes is the orthogonal problem we defer to follow-on work. A real product bug surfaced and was fixed during the study (a case/punctuation-insensitive value comparison), with the moat property (deterministic-supersession accuracy on clean code mutations) preserved and verified.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑