arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentMemBench:用于评估对话AI智能体长期记忆管理策略的系统基准

AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents

Ahmed Cherif

arXiv 2608.00009首次发表:更新:

发表机构

Sofrecom(索弗雷科姆)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究构建了AgentMemBench基准,评估五种记忆策略及两款现有系统,发现外部键值存储(EKV)在长程对话记忆任务中表现最优,但存在内存占用成本,同时公开了全部可复现资源。

AI 中文摘要

长期记忆仍是对话AI智能体的关键瓶颈,其有限的上下文窗口无法支持数千轮对话的连贯回忆。我们提出AgentMemBench,这是一个统一、可复现的基准,在相同条件下评估五种记忆管理策略:上下文窗口(ICW)、外部键值存储(EKV)、基于图的情景记忆(GEM)、基于压缩的摘要(CBS)以及基于网络的记忆(WAM)。所有策略均在三个公开数据集上进行评估,涵盖长期多轮对话(LoCoMo)、面向任务的文档接地(MultiDoc2Dial)和角色设定的多轮聊天(MSC),评估指标包括Recall@k、MRR、nDCG@k、答案F1、大语言模型评判的忠实度得分、内存占用以及491个带注释问题轮次的延迟。生成与评判均使用Qwen2.5-7B-Instruct(4-bit),采用贪婪解码以保证确定性。我们的结果显示:(1)EKV在所有质量维度上均占优(宏观Recall@5为0.792,MRR为0.677,F1为0.156,忠实度为0.354);(2)长程回忆起决定性作用:在黄金轮次位于多个会话之前的LoCoMo数据集上,ICW、WAM、GEM和CBS的Recall@5几乎为零(≤0.005),而仅EKV达到0.573,表明近期窗口、摘要和实体图在长程场景下失效,仅密集检索可扩展;(3)CBS在检索方面位列第二(0.556);(4)WAM因构造原因在语料内召回上与ICW相当,因为外部结果无语料内来源;(5)EKV的召回优势伴随内存占用成本(约5100个token,而ICW/WAM约为300个token),存在明确的准确率-效率权衡。我们还针对同一评估框架评估了两个已发布的记忆系统(MemGPT/Letta、HippoRAG),并发布了所有代码、环境和结果制品以实现完全可复现。

英文摘要

Long-term memory remains a critical bottleneck for conversational AI agents, whose finite context windows cannot support coherent recall across thousands of turns. We present AgentMemBench, a unified, reproducible benchmark evaluating five memory management strategies under identical conditions: in-context windowing (ICW), external key-value store (EKV), graph-based episodic memory (GEM), compression-based summarisation (CBS), and web-augmented memory (WAM). All are assessed across three public datasets covering long-term multi-session dialogue (LoCoMo), task-oriented document grounding (MultiDoc2Dial), and persona-grounded multi-session chat (MSC), using Recall@k, MRR, nDCG@k, Answer F1, an LLM-judge Faithfulness score, Memory Footprint, and Latency over 491 annotated question turns. Generation and judging both use Qwen2.5-7B-Instruct (4-bit), with greedy decoding for determinism. Our results show that (1) EKV dominates on every quality axis (macro Recall@5 0.792, MRR 0.677, F1 0.156, Faithfulness 0.354); (2) long-range recall is decisive: on LoCoMo, where the gold turn lies many sessions back, ICW, WAM, GEM, and CBS retrieve almost nothing (Recall@5 <= 0.005) while EKV alone reaches 0.573, showing that recency windows, summaries, and entity graphs collapse at long horizons and only dense retrieval scales; (3) CBS is the runner-up on retrieval (0.556); (4) WAM equals ICW on in-corpus recall by construction, since external results carry no in-corpus provenance; and (5) EKV's recall advantage carries a footprint cost (~5,100 vs ~300 tokens for ICW/WAM), an explicit accuracy-efficiency trade-off. We additionally evaluate two published memory systems (MemGPT/Letta, HippoRAG) against the same harness, and release all code, environment, and result artefacts for full reproducibility.

Comments22 pages, 3 figures submitted on Neural Computing and Applications

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑