arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VibeMemBench:评估编码智能体在真实仓库编码任务上的记忆系统

VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

Liyang Fan, Yingcheng Shi, Yongbin Li, Chenghao Sun, Xin Chen, Xander Xu, Hu Wei, Shiwen Ni, Min Yang, Jieping Ye

arXiv 2609.23570首次发表:更新:

发表机构

Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; SUAT; Alibaba Token Hub, Alibaba Group; Alibaba Group(中国科学院深圳先进技术研究院; SUAT; 阿里巴巴集团阿里巴巴令牌中心; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有评估未隔离记忆对编码结果影响的问题,提出VibeMemBench基准,基于90个仓库111个目标验证记忆系统,发现现有系统多数未能超越无记忆基线,揭示经验传递差距。

AI 中文摘要

编码智能体在真实仓库编码任务上运行,持久化记忆系统有望跨任务复用经验。然而,现有评估并未表明这些系统能否改善可执行的仓库工作。仓库基准测试代码变更但不隔离记忆,而记忆基准则对召回率打分,不衡量下游编码结果。我们提出VibeMemBench,一个用于评估记忆系统的基准,涵盖来自90个SWE-rebench V2仓库的111个编码目标,以及来自目标仓库的3,634条历史轨迹。这些目标遵循SWE基准风格,涵盖缺陷修复、功能请求、接口变更和配置工作。智能体在声明的记忆条件下编辑每个目标代码库,可执行测试决定任务是否解决。每个目标仅在注入的历史经验在参考设置中改善其可执行结果时保留,因此每个目标都携带先前经验,其有用性通过该设置中的执行得到验证。冻结的验证经验随后被转移到五个保留求解器。直接注入使其中四个求解器的观察任务解决率提高了1.1至4.5个百分点,同时降低了所有五个求解器的智能体步骤。然而,当四个现有记忆系统必须从相同历史中构建和检索经验时,十二个求解器与系统配对中有十一个未能超过匹配的无记忆基线。VibeMemBench揭示了仓库历史所持有的有用经验与现有记忆系统为仓库编码任务提供的经验之间的差距。

英文摘要

Coding agents operate on real repository coding tasks, and persistent memory systems promise to reuse experience across tasks. Yet existing evaluations do not show whether those systems improve executable repository work. Repository benchmarks test code changes but do not isolate memory, while memory benchmarks score recall without measuring downstream coding outcomes. We introduce VibeMemBench, a benchmark for evaluating memory systems on 111 coding targets from 90 SWE-rebench V2 repositories and 3,634 history trajectories from the target repositories. The targets follow the SWE benchmark style and cover bug fixes, feature requests, interface changes, and configuration work. An agent edits each target codebase under a declared memory condition. Executable tests decide task resolution. Each target is retained only when injected history experience improves its executable outcome in a reference setting, so every target carries a prior experience whose usefulness is verified by execution in that setting. The frozen verified experience is then transferred to five held-out solvers. Direct injection raises observed task resolution on four of them by 1.1 to 4.5 percentage points while lowering agent steps on all five. Yet when four existing memory systems must construct and retrieve experience from the same history, eleven of twelve solver and system pairings fail to exceed the matched memory-off baseline. VibeMemBench exposes the gap between the useful experience that repository history holds and the experience existing memory systems deliver for repository coding tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑