EngramEdit:通过条件记忆实现大语言模型中的解耦知识更新
EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory
查看机构详情
- The Hong Kong Polytechnic University(香港理工大学)
- Hangzhou Diagens Biotechnology Co., Ltd(杭州迪安基因生物技术有限公司)
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
EngramEdit通过条件记忆架构实现大语言模型的事实知识解耦更新,联合调整共享n-gram嵌入,在保持无关知识的同时达到近乎完美的编辑成功率,并支持多跳推理。
中文摘要 AI 辅助
诸如DeepSeek Engram之类的条件记忆架构利用输入n-gram来查找学习到的嵌入,以有限的额外计算扩展大语言模型(LLM)的能力。除了模型扩展之外,该架构还展示了将事实知识存储与通用计算解耦的潜力,为在保持Transformer主干固定的同时更新事实知识提供了一条有前景的途径。实现这一潜力具有挑战性,因为一个事实的不同表达可能激活不同的n-gram嵌入,而更新共享嵌入可能无意中改变模型对其他事实的预测。我们提出了EngramEdit,用于通过条件记忆进行解耦知识更新。EngramEdit首先计算目标记忆表示,使模型能够在多种表达中预测更新后的事实。然后,它联合更新共享的n-gram嵌入以匹配这些跨表达和编辑的目标,对频繁复用的嵌入的更新施加更强的惩罚,以保留无关知识。实验表明,EngramEdit能够通过条件记忆实现独立的事实知识更新,达到近乎完美的编辑成功率。修订后的知识可在未见过的表达和多跳推理中使用,在链式思维(CoT)提示下,其准确率几乎是强基线的三倍。即使事实更新累积,无关知识和通用能力也基本得到保留。这些发现表明,EngramEdit将条件记忆转变为可编辑的知识接口,将其作用扩展到模型扩展之外,以支持解耦的知识更新。
英文摘要
Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings can unintentionally change the model's predictions about other facts. We propose EngramEdit for decoupled knowledge updates through conditional memory. EngramEdit first computes target memory representations that make the model predict the updated fact across multiple expressions. It then jointly updates the shared n-gram embeddings to match these targets across expressions and edits, penalizing updates to frequently reused embeddings more strongly to preserve unrelated knowledge. Experiments show that EngramEdit enables independent factual knowledge updates through conditional memory, achieving near-perfect editing success. Revised knowledge is usable across unseen expressions and in multi-hop reasoning, with nearly three times the strongest baseline's accuracy under chain-of-thought (CoT) prompting. Unrelated knowledge and general capabilities are largely preserved even as factual updates accumulate. These findings show that EngramEdit turns conditional memory into an editable knowledge interface, extending its role beyond model scaling to support decoupled knowledge updates.