arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Raven:通过稀疏内存路由实现高召回率序列建模

Raven: High-Recall Sequence Modeling with Sparse Memory Routing

Arshia Afzal, Aviv Bick, Eric P. Xing, Volkan Cevher, Albert Gu

arXiv 2607.25357首次发表:更新:

发表机构

EPFL; Carnegie Mellon University; MBZUAI; Cartesia AI(洛桑联邦理工学院; 卡内基梅隆大学; 穆罕默德·本·扎耶德人工智能大学; 卡泰西亚人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对线性时间序列模型长上下文召回权衡问题,提出Raven模型,通过固定内存插槽和依赖输入的路由,减轻SWA和SSMs的不足,在召回密集基准测试中表现出色,外推时也保持有效。

AI 中文摘要

线性时间序列模型中的长上下文召回凸显了它们在写入内存方式上的权衡。基于状态的线性模型,如状态空间模型(SSMs)和线性Transformer,密集写入,每个新到达的令牌都会更新整个状态,这会导致干扰,使特定的过去令牌难以恢复。滑动窗口注意力(SWA)则相反,它通过存储显式令牌表示进行稀疏写入,但仅在固定窗口内,一旦相关令牌被逐出,召回率就会下降。在这些模型之间进行插值,我们引入了Raven,一个线性时间序列模型,它维护一组固定的内存插槽,并在每一步通过学习的、依赖输入的路由衰减和更新选定的子集。这使得Raven减轻了SWA基于位置的覆盖和硬逐出,同时减少了SSMs中密集状态更新的干扰,从而更有效地保留远程内容。在召回密集型基准测试中,Raven与之前的线性时间基线具有竞争力或表现更优,在SWA和SSMs急剧下降的情况下实现了强大的长上下文召回。当外推到高达其训练长度16倍的上下文长度时,它仍然有效,在混合架构中也有类似的提升。

英文摘要

Long-context recall in linear-time sequence models highlights a tradeoff in how they write to memory. State-based linear models, such as state-space models (SSMs) and linear Transformers, write densely, updating the entire state for each newly arrived token, which leads to interference and makes specific past tokens hard to recover. Sliding-window attention (SWA) exhibits the opposite behavior: it writes sparsely by storing explicit token representations, but only within a fixed window, so recall drops once the relevant token is evicted. Interpolating between these models, we introduce Raven, a linear-time sequence model that maintains a fixed set of memory slots and, at each step, decays and updates only a selected subset via learned, input-dependent routing. This lets Raven mitigate SWA's position-based overwriting and hard eviction while reducing interference from dense state updates in SSMs, thereby preserving long-range content much more effectively. Across recall-intensive benchmarks, Raven is competitive with or outperforms prior linear-time baselines, achieving strong long-context recall where both SWA and SSMs sharply degrade. It remains effective when extrapolating to context lengths as large as 16x its training length, with similar gains in hybrid architectures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑