发表机构
MemoryLake Team(MemoryLake团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究在MemoryArena的五个任务领域中,对比MemoryLake等三种智能体记忆后端,发现MemoryLake在数学等三个领域成功率最高,整体平均成功率领先,且性能与工作负载相关。
AI 中文摘要
大多数智能体记忆基准测试事后回忆能力,而MemoryArena评估记忆是否支持相互关联的多会话任务完成。我们在MemoryArena的全部五个领域中,对比了结构化多轨道记忆后端MemoryLake、文本嵌入模型text-embedding-3-small的向量RAG以及长上下文对照这三者的表现。这些系统采用相同的智能体框架、指定的gpt-5-mini模型别名、任务样本和评分代码,唯一刻意变更的组件是记忆集成模块。由于每个后端都整合了写入、检索、整合、预算管理和提示组装的相关选择,本研究是系统层面的匹配对比,而非仅针对表征的消融实验或成本匹配实验。在共享评估集上,MemoryLake在数学领域的观测成功率(SR)为9/40、物理领域为12/20、渐进式检索领域为4/20,均为最高;所有系统在旅行规划领域的SR均为0,网络购物领域仅长上下文系统实现了1次捆绑级成功(1/150),MemoryLake在旅行软过程得分和购物步骤匹配两项上均排名第三。遵循MemoryArena套件层面的惯例,对五个SR的事后等权重平均显示,MemoryLake为20.5%,表现最佳的对比系统为13.6%。这些为点估计:样本量较小,置信区间存在重叠,且未报告配对显著性检验。针对全部221个渐进式查询的仅MemoryLake单独运行,其计入失败的SR为26.7%(59/221),该结果不用于基线对比。研究结果支持记忆后端的性能取决于工作负载的观点,且在共享集上,四个被评估系统中MemoryLake表现领先;但该结果未确立基准范围内的最优水平,也未证明表征结构的因果优势。
英文摘要
Most agent-memory benchmarks test post-hoc recall, whereas MemoryArena evaluates whether memory supports interdependent, multi-session task completion. We compare MemoryLake, a structured multi-track memory backend, with Mem0, text-embedding-3-small vector RAG, and a long-context control across all five MemoryArena domains. The systems share the same agent framework, requested gpt-5-mini model alias, task samples, and scoring code; the memory integration is the intentionally changed component. Because each backend bundles write, retrieval, consolidation, budgeting, and prompt-assembly choices, the study is a matched system-level comparison, not a representation-only ablation or a cost-matched experiment. On the shared evaluation sets, MemoryLake has the highest observed success rate (SR) in mathematics (9/40), physics (12/20), and progressive retrieval (4/20). Every system has zero SR in travel planning, and web shopping yields a single bundle-level success (long context, 1/150); MemoryLake ranks third on both the travel soft process score and shopping step match. Following MemoryArena's suite-level convention, a post-hoc equal-weight average over the five SRs is 20.5% for MemoryLake versus 13.6% for the best comparator. These are point estimates: sample sizes are modest, confidence intervals overlap, and we do not report paired significance tests. A separate MemoryLake-only run over all 221 progressive queries yields a failure-counted SR of 26.7% (59/221) and is not a baseline comparison. The results support a workload-dependent view of memory backends and an observed lead among the four evaluated systems on the shared sets; they do not establish benchmark-wide state of the art or a causal advantage of representation structure.
Comments16 pages, 7 tables