发表机构
Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PRAGMA基准评估长期对话中个性化指导,发现现有系统在证据检索和利用上不足,需改进记忆架构。
AI 中文摘要
大型语言模型(LLMs)正越来越多地被部署为个性化助手,在较长时间内与用户互动。随着对话变得越来越长,依赖完整的交互历史变得越来越低效和不可靠:长上下文会引入大量的计算开销,使得模型难以一致地识别和利用当前请求最相关的信息。这些挑战促使了记忆系统的发展,这些系统能够构建和检索用户特定信息。在现实交互中,用户经常寻求实用指导,如推荐、规划和决策支持。与事实回忆任务不同,个性化指导要求模型整合跨多个过去对话的信息,并推理用户不断变化的偏好和体验。然而,现有的对话记忆评估主要侧重于检索和事实回忆。为了研究这一挑战,我们引入了PRAGMA,一个用于评估长期对话中个性化指导的基准。PRAGMA包含精心策划的纵向对话历史、证据注释和基于不断演变的用户上下文及错误用户假设的指导场景。跨检索系统、记忆系统和长上下文模型的实验表明,当前系统既难以恢复适当的对话证据,也难以有效地将其用于个性化指导。我们的结果强调了需要支持稳健的对话检索和超越证据回忆的基于记忆推理的记忆架构。
英文摘要
Large language models (LLMs) are increasingly deployed as personalized assistants that interact with users over extended periods of time. As conversations grow longer, relying on full interaction histories becomes increasingly inefficient and unreliable: long contexts introduce substantial computational overhead, making it difficult for models to consistently identify and utilize the most relevant information for the current request. These challenges have motivated memory systems that structure and retrieve user-specific information. In realistic interactions, users often seek practical guidance such as recommendations, planning, and decision support. Unlike factual recall tasks, personalized guidance requires models to integrate information across multiple past conversations and reason about changing user preferences and experiences. However, existing conversational memory evaluations mainly focus on retrieval and factual recall. To study this challenge, we introduce PRAGMA, a benchmark for evaluating personalized guidance in long-term conversations. PRGAMA contains curated longitudinal conversation histories, evidence annotations, and guidance scenarios grounded in evolving user contexts and incorrect user assumptions. Experiments across retrieval systems, memory systems, and long-context models reveal that current systems struggle both to recover the appropriate conversational evidence and to effectively use it for personalized guidance. Our results highlight the need for memory architectures that support robust conversational retrieval and memory-grounded reasoning beyond evidence recall.
CommentsAccepted to EMNLP 2026. v2: minor corrections to the RAG variant results in Table 5; main findings unchanged