MemUse:将记忆评估从直接问答转向人机长期对话中的自然整合
MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation
查看机构详情
- Graduate School of Informatics, Kyoto University(京都大学信息学研究科)
- SAP(思爱普)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究针对对话式LLM记忆系统评估问题,提出MemUse方法,发现直接问答准确率与用户满意度无关,而自然整合能力与之相关,同时发布了相关语料库、MemUse及配套资源。
中文摘要 AI 辅助
对话式大语言模型(LLM)的记忆系统通常通过针对先前对话的直接事实性问题(直接问答,Direct QA)进行评估,即模型能否从先前对话中回忆起事实X。我们在一项为期4个月的部署研究中测试了直接问答准确率是否与用户满意度相关,该研究包含40名用户、1872个会话和7种记忆条件。在这7种条件下,现有基准直接问答的准确率从19.7%到70.1%不等,但用户满意度并未发生变化。我们假设现有基准和用户满意度衡量的是不同的能力:基准衡量的是触发式检索(被询问时的回忆),而对话需要的是自然整合(检测相关性并将先前语境自然融入回复)。为验证这一点,我们引入了MemUse,这是一组从部署中提取的真实用户触发的记忆时刻,通过对自然对话回复的整合感知判断进行评分。在保持模型和语境固定的情况下,同一系统在直接问答上的得分为78.8%,但在对话中仅引用了这些事实的7.9%,存在71个百分点的差距。在这些时刻中,自然整合与满意度相关,而直接问答则不相关。我们在该https网址发布了部署语料库和MemUse,以及所有判断和评分提示。
英文摘要
Memory systems for conversational LLMs are conventionally evaluated by direct, fact-seeking questions about prior dialogue (Direct QA): can the model recall fact X from a prior conversation? We tested whether higher Direct QA accuracy correlates with higher user satisfaction in a 4-month deployment (40 users, 1,872 sessions, 7 memory conditions). Existing-benchmark Direct QA varies from 19.7% to 70.1% across the 7 conditions, but satisfaction does not change. We hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response). To examine this, we introduce MemUse, a set of real user-cued memory moments drawn from the deployment, scored by an integration-aware judgment of the natural conversational response. Holding the model and context fixed, the same system that scores 78.8% on Direct QA references only 7.9% of those facts in conversation -- a 71-point gap. Within these moments, Natural Integration is associated with satisfaction, whereas Direct QA is not. We release the deployment corpus and MemUse together with all judgments and scoring prompts at https://github.com/ryuichi-sumida/memuse.