arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当用户不提问时:对话智能体中上下文驱动的记忆检索基准测试

When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

Wen-Yu Chang, Yun-Nung Chen

arXiv 2609.03467首次发表:更新:

发表机构

National Taiwan University(台湾大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对对话智能体的记忆检索,构建了LOCOMO-CONV基准,发现对话式查询会暴露问答基准未察觉的检索差距,且检索能力与响应质量不完全相关,为记忆系统优化提供了方向。

AI 中文摘要

大型语言模型(LLMs)正越来越多地被用作长程对话智能体,这激发了人们对记忆系统日益增长的兴趣。然而,现有的基准测试主要通过问答式探测来评估记忆,而非原位对话使用。我们推出了LOCOMO-CONV,这是一个源自LOCOMO的对话记忆基准测试,包含四种查询类型:对话式、隐式、反事实和复合式。我们在五个代表性记忆系统上评估了检索召回率和端到端响应质量。实验表明,对话式框架揭示了问答式基准测试所忽视的巨大检索差距,尤其是在隐式和复合式查询上,多方面查询重写可缩小原始回合记忆的差距,但无法缩小抽象记忆的差距。我们进一步发现,强大的检索能力并不能完全转化为响应质量,且隐式查询存在隐性接地现象,即记忆可改善上下文接地,而无需明确呈现黄金事实。这些结果表明,基于推理的记忆细化是一个有前景的方向,我们还发布了辅助支持记忆注释,用于捕捉原始黄金证据之外的对话有用上下文。

英文摘要

Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-situ conversational usage. We introduce LOCOMO-CONV, a conversa- tional memory benchmark derived from Lo- CoMo with four query styles: dialog, implicit, counterfactual, and composed. Across five rep- resentative memory systems, we evaluate both retrieval recall and end-to-end response qual- ity. Our experiments show that conversational framing exposes substantial retrieval gaps over- looked by QA benchmarks, especially on im- plicit and composed queries, which multi-facet query rewriting narrows for raw-turn mem- ory but not abstractive memory. We further find that strong retrieval does not fully trans- late into response quality, and that implicit queries exhibit silent grounding, where mem- ory improves contextual grounding without ex- plicitly surfacing the gold fact. These results point to reasoning-based memory elaboration as a promising direction, and we release aux- iliary supportive_memory annotations captur- ing conversationally useful context beyond the original gold evidence.

CommentsAccepted by EMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑