发表机构
Shanghai Jiao Tong University; MemTensor (Shanghai) Technology Co., Ltd.; Theseus Lab(上海交通大学; MemTensor(上海)科技有限公司; 忒修斯实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MemTrace,一种来源感知的记忆系统,通过不可变记忆轨迹和记忆轨迹图保留执行历史,并在上下文受限时重建一致状态、验证证据有效性,在三个长时程编码基准上显著提升性能。
AI 中文摘要
随着编码智能体承担跨多个文件和阶段的长时程软件演化任务,更长的执行轨迹引入了两个相互关联的挑战:(1)累积的历史记录使上下文预算紧张,(2)仓库变更可能使早期的执行证据失效。现有方法通过扩大上下文窗口、压缩、检索或仓库表示等技术来应对这些挑战,但往往在上下文刷新后无法重建一致的任务状态,或无法验证所召回的证据是否仍然有效。为此,我们提出了MemTrace,一种具有来源感知的记忆系统,它保留执行历史,并使其重用与不断演化的任务(例如,迭代式跨文件修复)和仓库状态保持一致。MemTrace将历史存储为不可变的记忆轨迹,锚定到关键信息(如文件、符号、测试),并在记忆轨迹图中组织其执行顺序和依赖关系。当上下文受限时,工作记忆仅保留紧凑的记忆锚点,智能体可据此重建最新的执行状态并定位与其下一步行动相关的证据。在恢复历史证据之前,MemTrace会根据当前仓库状态检查其有效性,并且只检索下一步行动所需的内容。在三个互补的长时程编码基准上,MemTrace在相同骨干和测试框架下持续优于所有完全评估的基线,在Codex CLI下将DeepSWE pass@1提高了21.2个百分点,SWE-EVO解决率提高了4.4个百分点,SWE-Milestone分数提高了17.8个百分点。
英文摘要
As coding agents take on long-horizon software evolution tasks spanning multiple files and stages, longer execution trajectories introduce two coupled challenges: (1) accumulated histories strain context budgets, and (2) repository changes can invalidate earlier execution evidence. Existing approaches address these challenges through techniques like larger context windows, compression, retrieval, or repository representations, but often fail to reconstruct a consistent task state after a context refresh or verify whether recalled evidence remains valid. Thus, we introduce MemTrace, a provenance-aware memory system that preserves execution history and aligns its reuse with the evolving task (e.g., iterative cross-file repair) and repository state. MemTrace stores history as immutable Memory Traces anchored to key information (e.g., files, symbols, tests), and organizes their execution order and dependencies in a Memory Trace Graph. When context is constrained, working memory retains only compact Memory Anchors, from which the agent can reconstruct the latest execution state and locate evidence relevant to its next action. Before restoring historical evidence, MemTrace checks its validity against the current repository state and retrieves only what the next action requires. Across three complementary long-horizon coding benchmarks, MemTrace consistently outperforms all fully evaluated baselines under the same backbone and harness, improving DeepSWE pass@1 by 21.2 points, SWE-EVO Resolved Rate by 4.4 points, and SWE-Milestone Score by 17.8 points under Codex CLI.
Comments20 pages, 7 figures. Code: https://github.com/Homy-Xu/MemTrace