arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MemOps:在长期对话中对生命周期内存操作进行基准测试

MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations

Xixuan Hao, Zeyu Zhang, Zehao Lin, Yihang Sun, Ziliang Guo, Xichong Zhang, Yuxuan Liang, Feiyu Xiong, Zhiyu Li

arXiv 2607.12893首次发表:更新:

发表机构

MemTensor (Shanghai) Technology Co., Ltd.; The Hong Kong University of Science and Technology (Guangzhou); Renmin University of China(墨天轮(上海)科技有限公司; 香港科技大学(广州); 中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对长期对话中内存操作评估,引入MemOps基准测试,将对话内存视为操作生命周期,通过可控管道嵌入操作并评估,揭示了当前系统内存操作的问题,推动评估从答案评分转向操作级诊断。

AI 中文摘要

长期记忆已成为基于大语言模型的智能体的一项基础能力。现有基准测试几乎只通过下游问答来评估内存,这是一种黑箱方式,混淆了内存失败的多种原因。本文认为在动态长期交互中,内存是一系列显式操作的生命周期。我们引入MemOps基准测试,将对话内存重新表述为生命周期操作序列,通过可控生成管道将操作嵌入对话,在多种设置下评估。结果表明当前系统远非一致可靠,此研究将长期内存评估从最终答案评分转向可解释的操作级诊断。

英文摘要

Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, multi-session interactions. Existing benchmarks, however, evaluate such memory almost exclusively through downstream question answering, scoring only the correctness of a final answer. This black-box formulation conflates the heterogeneous causes of memory failure, such as missing the introduction of a relevant fact, binding an operation to the wrong target, or relying on stale values after a correction. As a result, it can credit correct answers despite their reliance on inconsistent or unsafe memory states. In this paper, we argue that, in dynamic long-horizon interactions, memory is not a static collection of facts but a lifecycle of explicit operations, including remembering, forgetting, updating, reflecting, and their compositions. We introduce MemOps, a benchmark that reformulates conversational memory as a sequence of lifecycle operations and represents each memory event with a structured trace specifying its trigger, target, scope, state transition, and supporting evidence. A controllable generation pipeline embeds these operations into long, task-oriented conversations and produces gold operation traces together with six categories of operation-level probes, evaluated under both adjacent-evidence and long-context settings. Across long-context, retrieval-based, parametric and managed-memory systems, MemOps disentangles failure modes that final-answer accuracy alone conceals, revealing that current systems remain far from uniformly reliable. For instance, session-level retrieval outperforms turn-level retrieval, and long-context models remain notably weak at reconstructing ordered memory-state trajectories. These results move long-term memory evaluation from final-answer scoring toward interpretable, operation-level diagnosis.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑