发表机构
Dartmouth College; Wuhan University(达特茅斯学院; 武汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出了用于评估对话智能体记忆利用能力的基准UtilMem,发现传统事实记忆基准表现优异的系统,其记忆利用能力未必出色,仅靠检索无法实现有效的记忆利用。
AI 中文摘要
长期记忆对于对话智能体的重要性日益提升,但现有基准主要通过逐点事实召回衡量记忆:即系统能否从过往交互中恢复孤立事实或事件级细节。然而,现实中的记忆利用往往需要更严苛的能力:将扩展交互历史中分散、隐含且带有噪声的证据整合为连贯、面向任务的输出,我们将这种能力称为记忆利用。本文引入UtilMem,这一诊断基准包含五个领域的1717个实例,旨在评估记忆利用的四个未被充分探索的方面:对密集历史的推理、识别隐含相关记忆、将分散证据合成为摘要、分析或计划,以及抵抗语义相似干扰项的干扰。通过评估一系列多样化的基于检索和记忆增强的系统,我们发现,在传统事实记忆基准上的优异表现并不能可靠地转化为有效的记忆利用。此外,仅靠检索是不够的:即使成功恢复了相关证据,系统也常常无法跨会话整合信息,或无法区分有用证据与看似合理的干扰项。这些发现揭示了访问存储信息与有效利用信息之间存在的巨大差距,并表明长期对话记忆的进展将需要明确支持证据整合和抵抗检索干扰的架构。代码可在该https URL获取。
英文摘要
Long-term memory is increasingly important for conversational agents, yet existing benchmarks primarily measure memory through pointwise factual recall: whether a system can recover isolated facts or event-level details from prior interactions. Real-world memory use, however, often requires a more demanding capability: integrating distributed, implicit, and noisy evidence across extended interaction histories into coherent, task-oriented outputs. We call this capability memory utilization. Here, we introduce UtilMem, a diagnostic benchmark comprising 1,717 instances across five domains, designed to evaluate four underexplored aspects of memory utilization: reasoning over dense histories, identifying implicitly relevant memories, synthesizing distributed evidence into summaries, analyses, or plans, and resisting interference from semantically similar distractors. Evaluating a diverse set of retrieval-based and memory-augmented systems, we find that strong performance on conventional factual-memory benchmarks does not reliably translate into effective memory utilization. Moreover, retrieval alone is insufficient: even when relevant evidence is successfully recovered, systems frequently fail to integrate information across sessions or to distinguish useful evidence from plausible distractors. These findings expose a substantial gap between accessing stored information and using it effectively, and suggest that progress in long-term conversational memory will require architectures that explicitly support evidence integration and robustness to retrieval interference. Code is available at https://github.com/peijunallin/UtilMem.