发表机构
FusionBrain Lab; MBZUAI; Cognitive AI Systems Lab; RUDN; London Institute for Mathematical Sciences; MIRAI; Lomonosov Moscow State University; Laboratory for Analysis and Controllable Text Generation Technologies RAS; School of Natural Sciences, Institute for Advanced Study, Princeton; Department of Brain Sciences, Weizmann Institute of Science(融合大脑实验室; 穆罕默德·本·扎耶德人工智能大学; 认知人工智能系统实验室; 俄罗斯人民友谊大学; 伦敦数学科学研究所; 未来人工智能研究机构; 莫斯科国立罗蒙诺索夫大学; 俄罗斯科学院分析与可控文本生成技术实验室; 普林斯顿高等研究院自然科学学院; 魏茨曼科学研究所脑科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何通过关联循环记忆扩展大语言模型上下文长度。提出ARMT方法,构建特定数据集,给出综合训练方法。实验表明该方法能让模型处理超长输入,更好泛化,且降低计算量。
AI 中文摘要
扩展大语言模型(LLMs)的上下文长度对许多实际应用至关重要,但标准变压器仍受二次计算和线性内存扩展的限制。本文研究关联循环记忆变压器(ARMT)作为在LLMs中实现长上下文处理、恒定内存扩展和更高效率的实用方法。主要贡献有:构建两个特定领域长上下文数据集评估实际工作负载;提出基于ARMT的上下文扩展综合训练方法,包括持续预训练等;实验表明ARMT增强模型能处理超长输入且性能不劣化,能更好泛化到分布外上下文长度,且在保持原上下文窗口内基线性能时所需FLOP少30%。
英文摘要
Extending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain constrained by quadratic compute and linear memory scaling. In this work, we investigate the Associative Recurrent Memory Transformer (ARMT) as a practical approach for enabling long-context processing in LLMs, constant memory scaling, and better efficiency. We make three main contributions. First, we construct two domain-specific long-context datasets designed to evaluate realistic workloads, focusing on narrow-domain fine-tuning scenarios. Second, we propose a comprehensive training recipe for ARMT-based context extension, combining continued pre-training, synthetic long-context data generation, curriculum learning, and selective integration of associative memory into chosen model layers. Third, we present an extensive experimental study demonstrating that ARMT-augmented models: (i) process inputs well beyond their original context limits without degrading performance relative to in-limit baselines; (ii) generalize more effectively to out-of-distribution context lengths; and (iii) need 30% less FLOPs while preserving baseline performance within the original context window.