arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

规模化的记忆解码器:一种预训练的参数化长期记忆

Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

Rubin Wei, Jiaqi Cao, Jiarui Wang, Junming Zhang, Qipeng Guo, Bowen Zhou, Zhouhan Lin

arXiv 2607.27919首次发表:更新:

AI 中文总结

本研究将记忆解码器扩展至69亿参数并在3000亿token上预训练,通过分布式Faiss等技术解决索引瓶颈,发现独立扩展记忆可更高效提升语言模型性能,在多基准测试中验证了其优势。

AI 中文摘要

仅解码器的语言模型将长期记忆与推理纠缠在单一参数集中,使得难以独立扩展记忆容量。记忆解码器(Memory Decoder)引入了参数化长期记忆模块,但仅在相对较小的规模下对其进行了研究。在本研究中,我们提出了规模化的记忆解码器,将记忆模型扩展至69亿参数,并在3000亿token上对其进行预训练。在该数据规模下,索引与搜索的综合成本使得标准的Faiss管道不可行。我们通过分布式Faiss索引与检索管道,结合稀疏的、按批次加载的kNN分布,解决了这一瓶颈问题。在不同模型规模下,我们发现与仅扩展基础模型相比,为记忆分配更多参数能获得更好的参数-性能权衡。在17个基准测试中,将69亿参数的通用记忆与Pythia-410M配对,使其平均得分从29.86提升至37.34,总参数比Pythia-12B(37.24)少39%却超越了它。对于06亿至140亿参数的Qwen3 Base模型,在每个规模下,17亿参数的领域记忆使三个领域的平均得分提升超过9个点。总体而言,我们的结果表明,独立扩展预训练记忆为提升语言模型性能提供了一种更具参数效率的路径。

英文摘要

Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens. At this data scale, the combined cost of indexing and search makes a standard Faiss pipeline infeasible. We address this bottleneck with a distributed pipeline for Faiss indexing and retrieval, together with sparse, batch-wise loading of kNN distributions. Across model scales, we find that allocating more parameters to memory yields a better parameter-performance tradeoff than scaling the base model alone. On 17 benchmarks, pairing a 6.9B general memory with Pythia-410M raises its average score from 29.86 to 37.34, surpassing Pythia-12B (37.24) with 39% fewer total parameters. For Qwen3 Base models ranging from 0.6B to 14B, 1.7B domain memories improve the average score across the three domains by more than 9 points at every scale. Overall, our results demonstrate that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑