arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22231cs.CL

EvalMem:长期记忆系统的操作级诊断框架

EvalMem: An Operation-Level Diagnostic Framework for Long-Term Memory Systems

  • Nankai University(南开大学)
  • vivo AI Lab(vivo AI实验室)
  • Tsinghua University(清华大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

Zeyu Liu, Jian Zhong, Rongduo Han, Ziyang Wu, Shunye Tang, Chenghao He, Yaxuan Yang, Yihang Qiu, Ailing Wang, Xiao Liang, Guohuan Xie, Xiaokang Xue, Gongchen Li… 展开作者

Zeyu Liu, Jian Zhong, Rongduo Han, Ziyang Wu, Shunye Tang, Chenghao He, Yaxuan Yang, Yihang Qiu, Ailing Wang, Xiao Liang, Guohuan Xie, Xiaokang Xue, Gongchen Li, Haining Zhang, Wei Wang

AI总结:

EvalMem通过三个并行检查器对长期记忆系统进行操作级诊断,定位检索为最常见失败层,并借助MemWiki辅助结构提升平均准确率2.5和2.3个百分点。

AI中文摘要:

与基于LLM的助手进行长时程交互需要能够保存和更新用户状态、偏好及交互历史的记忆系统。现有评估仅报告端到端的问答准确率,无法确定错误究竟源于编码、检索还是生成环节。我们提出EvalMem,一个操作级诊断框架,包含三个并行的检查器。对于每个查询,编码检查器检查目标事实是否被存储,检索检查器评估原生检索器是否返回可用证据,生成检查器测试模型能否仅凭oracle证据作答。它们的输出构成细粒度的多标签缺陷代码。为改进存储级诊断,我们采用自适应agentic RAG,并采用先召回策略,同时使用查询和源证据进行搜索,将LoCoMo上现有证据的召回率从70.2%提升至95.6%。在LoCoMo、LongMemEval-S和动态DynaMem-Bench上对七个记忆系统的评估表明,检索是最常被归因的失败层;在默认LoCoMo设置下,检索缺陷达22.1%,而编码缺陷为7.7%,生成缺陷为6.5%。在此诊断指导下,MemWiki(一种基于各系统记忆导出构建的、利于搜索的辅助结构)在LoCoMo和LongMemEval-S上分别将平均准确率提升了2.5和2.3个百分点。

英文摘要:

Long-horizon interactions with LLM-based assistants require memory systems that preserve and update user states, preferences, and interaction histories. Existing evaluations report end-to-end QA accuracy and cannot determine whether errors arise from encoding, retrieval, or generation. We introduce EvalMem, an operation-level diagnostic framework with three parallel Examiners. For each query, the Encoding Examiner checks whether the target fact is stored, the Retrieval Examiner assesses whether the native retriever returns usable evidence, and the Generation Examiner tests whether the model can answer from oracle evidence. Their outputs form fine-grained multi-label defect codes. To improve store-level diagnosis, we adapt agentic RAG with a recall-first strategy that searches using both the query and source evidence, increasing recall of present evidence on LoCoMo from 70.2% to 95.6%. Evaluations of seven memory systems on LoCoMo, LongMemEval-S, and dynamic DynaMem-Bench identify retrieval as the most frequently attributed failure layer; in default LoCoMo, retrieval defects reach 22.1%, compared with 7.7% for encoding and 6.5% for generation. Guided by this diagnosis, MemWiki, a search-friendly auxiliary structure built from each system's memory export, improves mean accuracy by 2.5 and 2.3 percentage points on LoCoMo and LongMemEval-S, respectively.

补充信息

↑