发表机构
UNC Chapel Hill; The University of Texas at Austin; Mila; Yale University(北卡罗来纳大学教堂山分校; 德克萨斯大学奥斯汀分校; 米拉研究所; 耶鲁大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发现混合语言模型过度依赖注意力而忽视循环记忆通路,提出辅助损失以限制注意力早期访问,促进循环通路使用,提升长上下文与信息聚合任务性能。
AI 中文摘要
循环-注意力混合语言模型(LMs)交错使用注意力层和循环层,日益被用于结合循环层的效率与注意力层的强性能。先前的研究表明,注意力层和循环层提供了利用过去信息的互补通路:注意力支持从较早的token进行精确记忆回忆,而循环层支持在长上下文中整合分散信息。然而,我们观察到,仅仅拥有这两条通路并不意味着混合语言模型正在有效使用它们。我们发现,它们对注意力的依赖远大于对循环状态的依赖。标准的监督微调提高了整体性能,但并未改善两条记忆通路的协调方式:模型变得更加依赖由注意力层传播的信息,而其对由循环层传播的信息的使用仍然有限。为了鼓励两条记忆通路之间更好的协调,我们添加了一个辅助损失,限制注意力对较早上下文的访问,同时循环状态在整个序列中传播。这一目标鼓励模型通过循环通路与注意力一起保留和使用信息。它提高了整体性能,特别是在涉及较长上下文或需要信息聚合的任务上取得了显著提升,这与分析中观察到的循环层的优势一致。至关重要的是,这种不平衡和我们辅助损失的益处具有普遍性:它们适用于问答和智能体任务中的多个循环-注意力语言模型,以及结合不同记忆形式的基于注意力的语言模型。总之,我们的发现表明,仅仅提供多条记忆通路并不能确保其有效使用,需要有针对性的监督来更好地协调它们。
英文摘要
Recurrent-attention hybrid language models (LMs), which interleave attention and recurrent layers, are increasingly used to combine the efficiency of the recurrent layers with the strong performance of attention layers. Prior work suggests that attention and recurrent layers offer complementary pathways to use past information: attention supports precise memory recall from earlier tokens, while recurrent layers support consolidation of disparate information over long contexts. However, we observe that simply having access to both pathways does not mean that hybrid LMs are effectively using them. We find that they rely substantially more on attention than on the recurrent state. Standard supervised fine-tuning improves overall performance but does not improve how the two memory pathways are coordinated: the model becomes more reliant on information propagated by attention layers, while its use of information propagated by recurrent layers remains limited. To encourage better coordination between the two memory pathways, we add an auxiliary loss that limits attention's access to earlier context while the recurrent state propagates through the full sequence. This objective encourages the model to retain and use information through the recurrent pathway alongside attention. It improves overall performance, with particularly strong gains on tasks involving longer contexts or requiring information aggregation, consistent with the strengths of recurrent layers observed in analysis. Crucially, this imbalance and the benefit of our auxiliary loss generalize: they apply to multiple recurrent-attention LMs in question-answering and agentic tasks, as well as to attention-based LMs that combine different forms of memory. Together, our findings show that simply providing multiple memory pathways does not ensure their effective use, and that targeted supervision is needed to better coordinate them.
CommentsCode: https://github.com/amy-hyunji/Balancing-Memory-Pathways