arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

记忆规范化:持久化LLM记忆中跨模型漂移的框架与基准

Memory Canonicalization: A Framework and Benchmark for Cross-Model Drift in Persistent LLM Memory

Amit Vadnere, Aishwarya Lonarkar

arXiv 2610.05124首次发表:更新:

发表机构

Google(谷歌)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出记忆规范化框架,通过写入时管线重写LLM持久记忆对象以消除歧义和情感负载,并构建CMSC-E基准,实验显示情感一致性有初步提升但未通过多重比较校正。

AI 中文摘要

大型语言模型(LLM)的持久化记忆已迅速成熟:诸如MemGPT/Letta、Mem0和Zep等系统现在为智能体提供分层、时间感知、模型无关的外部存储,而模型上下文协议(MCP)则标准化了对记忆服务器的访问。一个较少被关注的问题是:在完全相同的条件下,两个不同的LLM检索到的相同存储记忆对象,可能在事实或情感上被不同地解读。本文提出记忆规范化:一种写入时管线,用于检测原始记忆对象中的歧义、条件结构和情感负载,并将其重写为显式、结构上消除歧义的规范形式,其中情感效价作为单独字段表示,而非从语气中推断。我们形式化了该管线,定义了一个配套的跨模型语义漂移/情感一致性评分基准(CMSC-E),并报告了一项使用176个合成记忆对象和三个下游模型家族的三臂试验结果。我们发现,相对于原始记忆,完全规范化的记忆在跨模型情感一致性上存在未校正的改进(+0.050,95%自助法置信区间[0.013, 0.086],配对t检验p=0.010),但该结果在Bonferroni、Holm或Benjamini-Hochberg校正下,在测试的六项比较中均未存活。所有事实漂移(CMSD)比较在任何校正水平下均未达到显著性。我们将这些结果报告为探索性而非验证性,并概述了后续必要工作,包括更大样本、独立评判模型、人工验证的渲染和预注册。

英文摘要

Persistent memory for Large Language Models (LLMs) has matured rapidly: systems such as MemGPT/Letta, Mem0, and Zep now provide agents with tiered, temporally-aware, model-agnostic external storage, while the Model Context Protocol (MCP) standardizes access to memory servers. A less addressed problem is that an identical stored memory object, retrieved by two different LLMs under otherwise identical conditions, may not be interpreted the same way, factually or emotionally. This paper proposes memory canonicalization: a write-time pipeline that detects ambiguity, conditional structure, and emotional loading in a raw memory object and rewrites it into an explicit, structurally disambiguated canonical form, with emotional valence represented as a separate field rather than inferred from tone. We formalize the pipeline, define a companion Cross-Model Semantic Drift / Emotional Consistency Score benchmark (CMSC-E), and report results from a three-arm pilot using 176 synthetic memory objects and three downstream model families. We find an uncorrected improvement in cross-model emotional consistency for fully canonicalized memory relative to raw memory (+0.050, 95% bootstrap CI [0.013, 0.086], paired t-test p = 0.010), but this result does not survive Bonferroni, Holm, or Benjamini-Hochberg correction across the six comparisons tested. None of the factual-drift (CMSD) comparisons reach significance at any correction level. We report these results as exploratory rather than confirmatory and outline needed follow-up work, including larger samples, independent judge models, human-validated rendering, and preregistration.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑