发表机构
University of Toronto; Vector Institute; Université de Montréal; Mila - Quebec AI Institute; RIKEN AIP; Western University(多伦多大学; 向量研究所; 蒙特利尔大学; Mila-魁北克人工智能研究所; 日本理化学研究所革新智能统合研究中心; 韦仕敦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ARM,一种基于路由内存的KV缓存结构,通过Gumbel-Softmax和可学习策略实现动态内存选择,在长上下文推理中减少信息损失并提升性能与效率。
AI 中文摘要
尽管长上下文推理取得了进展,但大型语言模型(LLMs)仍从根本上受限于为稳定计算所必需的关键值(KV)缓存机制。诸如选择性令牌驱逐和剪枝等技术已大大缓解了这些问题,但为了管理不断增长的缓存,这些技术常常丢弃核心信息。在本文中,我们提出了基于路由内存的注意力机制(ARM),这是一种新颖的KV缓存结构,引入了一个完全可微的、固定大小的内存系统,该系统组织为分层路由器。通过Gumbel-Softmax,ARM学习选择内存槽并执行sigmoid门控更新,以柔和地组合新信息和存储信息,避免硬驱逐并减少信息损失。通过进一步训练一个策略,在推理时动态选择不同数量的内存,ARM能够针对简单上下文和需要更深层推理的输入调整其访问,从而在短上下文和长上下文上实现更可扩展和更有效的检索。在标准常识和长上下文推理基准上的实验结果表明,与固定KV缓存方法相比,ARM在性能和效率上均实现了优越的表现,同时在内存和生成延迟方面保持高效和可扩展。
英文摘要
Despite advances in long-context inference, large language models (LLMs) remain fundamentally limited by the key-value (KV) caching mechanisms that are necessary for stable computation. Techniques such as selective token eviction and pruning have vastly mitigated these issues, but often discard core information to manage the growing cache. In this paper, we propose Attention with Routed Memory (ARM) a novel KV caching structure that introduces a fully differentiable, fixed-size memory system organized as a hierarchical router. Via a Gumbel-Softmax, ARM learns to select memory slots and perform sigmoid-gated updates that softly combine new and stored information, avoiding hard eviction and reducing information loss. By further training a policy to dynamically select varying amounts of memory at inference, ARM adapts its accesses for both simple contexts and inputs that require deeper reasoning, enabling more scalable and effective retrieval on both short- and long-contexts. Experimental results on standard commonsense and long-context reasoning benchmarks demonstrate that ARM achieves superior performance and efficiency compared to fixed KV-caching approaches, while remaining efficient and scalable in terms of both memory and generation latency.
CommentsAccepted to the Forty-third International Conference on Machine Learning (ICML) 2026. First two authors contributed equally