arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基础的裂痕:看似微小的架构选择会影响长上下文扩展

Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension

Amanda Bertsch, Luca Soldaini, Matthew R. Gormley, Graham Neubig, Hannaneh Hajishirzi, Kyle Lo, Dirk Groeneveld

arXiv 2608.10296首次发表:更新:

发表机构

Ai2; Carnegie Mellon University; University of Washington(Ai2; 卡内基梅隆大学; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现,Olmo、Llama、Qwen等稠密模型系列的四个微小架构决策会复合降低长上下文性能,结合三个及以上选择可使性能降47%,经17万余GPU小时训练发布OlmPool,其部分模型长上下文性能优于Llama 3。

AI 中文摘要

人们可能会认为,在稠密Transformer范式内的架构变体对准确率的影响有限。然而,我们证明在长上下文场景下并非如此。具体而言,我们展示了四个微小的架构决策——这些决策至少被Olmo、Llama和Qwen稠密模型系列中的一个所采用——会对长上下文可扩展性产生复合负面影响。单独其中任何一个选择对长上下文性能的影响都很小,但结合三个或更多选择会使下游性能下降多达47%。此外,这些差异无法从短上下文损失或验证数据集中检测到。我们表明,不同模型系列之间长上下文能力的大部分差异是由这些架构特征驱动的,并且可以通过在预训练早期应用上下文扩展来检测。我们通过控制变量的消融实验证明了这一点,这些实验在固定数据、分词器和扩展方案的同时,改变归一化、GQA(分组查询注意力)、预训练上下文长度和滑动窗口注意力。经过超过170000 GPU小时的训练后,我们发布了由此产生的模型集合OlmPool,这是一组26个可比的7B模型,包含长上下文扩展前后的检查点。该集合包括几种在长上下文可扩展性上优于Llama 3架构的架构。在对我们的消融模型的分析中,我们确定了注意力汇聚行为和跨上下文注意力分布的模式,这些模式可归因于特定的架构差异。

英文摘要

One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the long context setting. Specifically, we show that a set of four minor architectural decisions --- all made by at least one of the Olmo, Llama, and Qwen dense model families --- have a compoundingly negative effect on long context extensibility. Any one of these choices alone has a minor impact on long context performance, but combining three or more can drop the performance downstream by up to 47%. Furthermore, these differences are not detectable from short-context loss or validation datasets. We show that much of the variation in long context ability across model families is driven by these architectural features and detectable from applying context extension early in pretraining. We demonstrate this with controlled ablations that hold data, tokenizer, and extension recipe fixed while varying normalization, GQA, pretraining context length, and sliding window attention. After over 170,000 GPU hours of training, we release the resulting set of models as OlmPool, a set of 26 comparable 7B models with checkpoints before and after long-context extension. This pool includes several architectures that outperform the Llama 3 architecture on long context extensibility. In an analysis of our ablation models, we identify patterns in attention sink behavior and attention distributions across context that are attributable to specific architectural differences.

Comments29 pages; accepted to COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑