arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过关联循环记忆扩展大语言模型上下文

Extending LLM Context via Associative Recurrent Memory

Gleb Kuzmin, Ivan Rodkin, Aydar Bulatov, Yuri Kuratov, Lyudmila Rvanova, Mikhail Katkov, Ilia Sochenkov, Misha Tsodyks, Timothy Baldwin, Mikhail Burtsev, Artem Shelmanov

arXiv 2607.11614首次发表:更新:

发表机构

FusionBrain Lab; MBZUAI; Cognitive AI Systems Lab; RUDN; London Institute for Mathematical Sciences; MIRAI; Lomonosov Moscow State University; Laboratory for Analysis and Controllable Text Generation Technologies RAS; School of Natural Sciences, Institute for Advanced Study, Princeton; Department of Brain Sciences, Weizmann Institute of Science(融合大脑实验室; 穆罕默德·本·扎耶德人工智能大学; 认知人工智能系统实验室; 俄罗斯人民友谊大学; 伦敦数学科学研究所; 未来人工智能研究机构; 莫斯科国立罗蒙诺索夫大学; 俄罗斯科学院分析与可控文本生成技术实验室; 普林斯顿高等研究院自然科学学院; 魏茨曼科学研究所脑科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何通过关联循环记忆扩展大语言模型上下文长度。提出ARMT方法,构建特定数据集,给出综合训练方法。实验表明该方法能让模型处理超长输入,更好泛化,且降低计算量。

AI 中文摘要

扩展大语言模型(LLMs)的上下文长度对许多实际应用至关重要,但标准变压器仍受二次计算和线性内存扩展的限制。本文研究关联循环记忆变压器(ARMT)作为在LLMs中实现长上下文处理、恒定内存扩展和更高效率的实用方法。主要贡献有:构建两个特定领域长上下文数据集评估实际工作负载;提出基于ARMT的上下文扩展综合训练方法,包括持续预训练等;实验表明ARMT增强模型能处理超长输入且性能不劣化,能更好泛化到分布外上下文长度,且在保持原上下文窗口内基线性能时所需FLOP少30%。

英文摘要

Extending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain constrained by quadratic compute and linear memory scaling. In this work, we investigate the Associative Recurrent Memory Transformer (ARMT) as a practical approach for enabling long-context processing in LLMs, constant memory scaling, and better efficiency. We make three main contributions. First, we construct two domain-specific long-context datasets designed to evaluate realistic workloads, focusing on narrow-domain fine-tuning scenarios. Second, we propose a comprehensive training recipe for ARMT-based context extension, combining continued pre-training, synthetic long-context data generation, curriculum learning, and selective integration of associative memory into chosen model layers. Third, we present an extensive experimental study demonstrating that ARMT-augmented models: (i) process inputs well beyond their original context limits without degrading performance relative to in-limit baselines; (ii) generalize more effectively to out-of-distribution context lengths; and (iii) need 30% less FLOPs while preserving baseline performance within the original context window.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑