你的智能体的记忆能在模型升级后保留吗?一项关于记忆可移植性的对照研究
Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability
- LinkedIn(领英)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究对照分析了LC-RAW、RAG、NOTES、KG-fixed四种记忆存储方式在模型升级后的可移植性,发现KG-fixed迁移可靠,NOTES和RAG存在显著性能损失,需重视迁移测试与源历史保留。
AI中文摘要:
模型升级是常规操作,但记忆迁移却并非如此。智能体即便保留相同的记忆存储仍可能出现遗忘:新模型可能对旧笔记的解读不同,混合嵌入版本可能破坏检索,且若无原始证据则修复可能失败。我们对记忆进行了比较,将相同历史原封不动地保留用于长上下文阅读(LC-RAW)、分块后用于检索增强生成(RAG)、由模型压缩为自然语言笔记(NOTES),或归一化为固定模式知识图谱(KG-fixed)。本研究使用48个带有随机答案编码的合成历史、精确评分,以及两个参数规模低于100亿的开放权重模型。我们的测量结果显示,固定模式结构的迁移表现可靠,KG-fixed的准确率在更换写入模型后仅变化+0.0004±0.0020;相反,压缩后的NOTES表现出高度的模型耦合性,根据具体迁移方向,准确率会不对称地变化+9.91或-13.28个百分点。在RAG系统中,使用50/50混合索引的部分嵌入迁移仅能实现4.96个百分点的准确率提升,而完全重新嵌入可获得11.90个百分点的提升,前者放弃了大部分增益。诊断分解结果显示,NOTES准确率缺陷的80%(0.467±0.014)源于初始构建过程中丢失的信息,而RAG缺陷的81%(0.364±0.012)由检索失败导致。最后,仅对NOTES进行存储修复在全部48个测试案例中均未达到90%的性能恢复目标,而保留原始源历史则在一个测试方向的48个案例中实现了34个的成功恢复。这些发现强调了针对具体迁移方向进行测试、严格隔离嵌入空间以及保留源历史用于记忆修复的必要性。
英文摘要:
Model upgrades are routine; memory migrations are not. An agent can keep the same memory store and still forget: a new model may interpret old notes differently, mixed embedding versions may break retrieval, and repair may fail without the original evidence. We compare memory as the same history is preserved verbatim for long-context reading (LC-RAW), divided into chunks for retrieval-augmented generation (RAG), compressed by a model into natural-language notes (NOTES), or normalized into a fixed-schema knowledge graph (KG-fixed). The study uses 48 synthetic histories with randomized answer codes, exact scoring, and two open-weight models with sub 10 billion parameters. Our measurements show that fixed-schema structures transfer reliably, with KG-fixed accuracy changing by only $+0.0004 \pm 0.0020$ following a writer swap. Conversely, compressed NOTES exhibit high model coupling, with accuracy shifting asymmetrically by $+9.91$ or $-13.28$ percentage points depending on the specific migration direction. In RAG systems, partial embedding migrations using a 50/50 mixed index capture only a 4.96-point accuracy improvement, forfeiting the majority of the 11.90-point gain achieved through full re-embedding. Diagnostic decomposition attributes 80% ($0.467 \pm 0.014$) of the NOTES accuracy deficit to information lost during initial construction, whereas retrieval failures drive 81% ($0.364 \pm 0.012$) of the RAG deficit. Finally, store-only repair of NOTES fails to reach a 90% performance recovery target in all 48 test cases, whereas retaining the raw source history enables successful recovery in 34 of 48 cases for one tested direction. These findings highlight the necessity of direction-specific migration testing, strict embedding space isolation, and the retention of source histories for memory repair.