arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DreamBench-SWE:面向软件智能体的多会话记忆卫生基准

DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents

Sarthak Singh

arXiv 2608.20664首次发表:更新:

AI 中文总结

该研究提出面向软件智能体的多会话记忆卫生基准DreamBench-SWE,通过v2和预注册的v2.1审计,对比不同内存配置的任务完成表现,验证其作为可执行基准的有效性,同时明确相关内存机制的表现情况。

AI 中文摘要

DreamBench-SWE是一个面向软件智能体记忆卫生的多会话基准,后续软件任务依赖于早期会话中无法推断的证据,并通过可执行的隐藏预言机进行评分。我们报告了原始的缩放v2折叠版本,以及一项独立预注册的v2.1后续审计,该审计在该研究之后设计,但在检查后续结果前已冻结。后续运行在四种条件下完成了360/360个工作单元和720/720个S3单元。在原始折叠版本中,主要的DF-hybrid--B5对比无效(95/180对89/180;聚类p=0.518,Holm校正后p=1),并非等价的证据,且C9/C10保留了B0的余量限制。在后续版本中,无外部内存实现21/180次通过(通过率0.1167),确定性逐字事件内存实现82/180次通过(通过率0.4556),类型化加原始参考探针实现83/180次通过(通过率0.4611),以及一个固定托管的Mem0文字存储配置实现97/180次通过(通过率0.5389)。预注册的六槽Family A在p=1时保留不可用槽;与无内存的所有三次可用对比在Holm校正后被拒绝。两项预注册的机制对比在预评估一致性拒绝后不可用。次要的文字存储与逐字对比非确证且依赖敏感性,而与参考探针的对比未被拒绝。因此,该审计支持DreamBench-SWE作为具有区分性的可执行轮廓基准,并表征了一个确切的托管内存配置,但未确立外部系统机制、带内存条件的优越性、等价性或广泛的产品通用性。原始v2.0.5的发现和产物保持不变。

英文摘要

DreamBench-SWE is a multi-session benchmark for software-agent memory hygiene in which later software tasks depend on non-inferable evidence from earlier sessions and are scored by executable hidden oracles. We report the original scaled v2 fold and a separately preregistered v2.1 successor audit designed after that study but frozen before successor outcome inspection. The successor run completed 360/360 work units and 720/720 S3 cells across four conditions. In the original fold, the primary DF-hybrid--B5 contrast was null (95/180 versus 89/180; clustered p=.518, Holm p=1), not evidence of equivalence, and C9/C10 retained B0-headroom limitations. In the successor, no external memory achieved 21/180 passes (rate 0.1167), deterministic verbatim event memory 82/180 (rate 0.4556), the typed-plus-raw reference probe 83/180 (rate 0.4611), and one pinned hosted Mem0 literal-storage configuration 97/180 (rate 0.5389). The registered six-slot Family A retained unavailable slots at p=1; all three available comparisons against no memory rejected after Holm correction. Both preregistered mechanism contrasts were unavailable after pre-evaluation conformance rejection. The secondary literal-storage-versus-verbatim comparison was nonconfirmatory and sensitivity-dependent, while the comparison with the reference probe did not reject. The audit therefore supports DreamBench-SWE as a discriminating executable profile benchmark and characterizes one exact hosted-memory configuration, but it does not establish an external-system mechanism, superiority among memory-bearing conditions, equivalence, or broad product generality. The original v2.0.5 findings and artifacts remain unchanged.

Comments67 pages. Public benchmark and evidence release: https://github.com/iroiro147/dreambench-swe/releases/tag/v2.1.0

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑