arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过目标侧读取器适配实现跨模型记忆迁移

Frozen Memory Is Not Enough: Rethinking External Memory as Extraction

Mingyuan Li, Guangsheng Yu, Xu Wang, Shaoxiong Ji

arXiv 2608.17050首次发表:更新:

AI 中文总结

该研究针对大语言模型知识利用的两类方法缺陷,探究Engram式哈希记忆跨模型迁移的关键因素,通过跨模型冻结记忆提取实验,提出双层四分支读取器,实现同模型与跨模型复用差距的缩小,证明Engram可作为可复用外部知识人工制品。

AI 中文摘要

提升大语言模型知识利用的方法通常分为两类:非参数检索可灵活访问外部知识,但会增加检索延迟、上下文开销,且仅与主干模型进行浅层集成;参数化适配在推理时效率较高,但会将知识与模型权重绑定,难以更新、审计或迁移。Engram式哈希记忆处于中间状态:它将学习到的信息存储在外部可寻址表中,仅通过小型学习读取器使用该表。这引发了一个基本问题:当此类记忆跨主干模型迁移时,更重要的是冻结的记忆本身还是目标侧读取器?我们通过跨模型冻结记忆提取研究该问题:将在源模型上训练的记忆冻结并附加到不同的目标模型,仅训练轻量读取器。消融实验显示,学习到的记忆内容和正确寻址都很重要,但迁移后的表仅在与目标模型对齐的读取器下才有用。在下游问答任务中,双层四分支读取器几乎缩小了同模型复用与跨模型复用之间的差距,在我们的受控评估协议下平均得分达38.8。此外,当提供的读取器与目标接口直接兼容时,冻结的人工制品无需目标侧训练即可提供显著效用,而可选的读取器适配可进一步提升性能。这些结果表明,只要目标具备兼容的读取器接口,Engram可作为可复用的外部知识人工制品;当直接读取器复用不足时,目标侧适配可进一步提升对齐效果。

英文摘要

Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer. Engram-style hashed memory occupies a middle regime: it stores learned information in an external, addressable table, yet consumes that table through a small learned reader. This raises a basic question: when such a memory is moved across backbones, what matters more, the frozen memory itself or the target-side reader? We study this question through cross-model frozen-memory extraction, in which a memory trained on a source model is frozen and attached to a different target model, with only a lightweight reader trained. Ablations show that learned memory content and correct addressing both matter, but the transferred table becomes useful only through a reader aligned to the target model. In downstream question answering tasks, a dual-layer, four-branch reader nearly closes the gap between same-model and cross-model reuse, achieving an average score of 38.8 under our controlled evaluation protocol. Moreover, when the provider reader is directly compatible with the target interface, the frozen artifact can provide substantial utility without target-side training, while optional reader adaptation yields further improvement. These results suggest that Engram can serve as a reusable external knowledge artifact, provided that the target has access to a compatible reader interface; target-side adaptation can further improve alignment when direct reader reuse is insufficient.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑