发表机构
MemTensor (Shanghai) Technology Co., Ltd.; The Hong Kong University of Science and Technology (Guangzhou); Renmin University of China(墨天轮(上海)科技有限公司; 香港科技大学(广州); 中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对长期对话中内存操作评估,引入MemOps基准测试,将对话内存视为操作生命周期,通过可控管道嵌入操作并评估,揭示了当前系统内存操作的问题,推动评估从答案评分转向操作级诊断。
AI 中文摘要
长期记忆已成为基于大语言模型的智能体的一项基础能力。现有基准测试几乎只通过下游问答来评估内存,这是一种黑箱方式,混淆了内存失败的多种原因。本文认为在动态长期交互中,内存是一系列显式操作的生命周期。我们引入MemOps基准测试,将对话内存重新表述为生命周期操作序列,通过可控生成管道将操作嵌入对话,在多种设置下评估。结果表明当前系统远非一致可靠,此研究将长期内存评估从最终答案评分转向可解释的操作级诊断。
英文摘要
Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, multi-session interactions. Existing benchmarks, however, evaluate such memory almost exclusively through downstream question answering, scoring only the correctness of a final answer. This black-box formulation conflates the heterogeneous causes of memory failure, such as missing the introduction of a relevant fact, binding an operation to the wrong target, or relying on stale values after a correction. As a result, it can credit correct answers despite their reliance on inconsistent or unsafe memory states. In this paper, we argue that, in dynamic long-horizon interactions, memory is not a static collection of facts but a lifecycle of explicit operations, including remembering, forgetting, updating, reflecting, and their compositions. We introduce MemOps, a benchmark that reformulates conversational memory as a sequence of lifecycle operations and represents each memory event with a structured trace specifying its trigger, target, scope, state transition, and supporting evidence. A controllable generation pipeline embeds these operations into long, task-oriented conversations and produces gold operation traces together with six categories of operation-level probes, evaluated under both adjacent-evidence and long-context settings. Across long-context, retrieval-based, parametric and managed-memory systems, MemOps disentangles failure modes that final-answer accuracy alone conceals, revealing that current systems remain far from uniformly reliable. For instance, session-level retrieval outperforms turn-level retrieval, and long-context models remain notably weak at reconstructing ordered memory-state trajectories. These results move long-term memory evaluation from final-answer scoring toward interpretable, operation-level diagnosis.