arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MARCH:通过内容路由状态锚点扩展循环记忆

MARCH: Scaling Recurrent Memory with Content-Routed State Anchors

Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Ning Ding, Xia Hu, Bowen Zhou, Chaochao Lu, Youbang Sun

arXiv 2608.12435首次发表:更新:

发表机构

Shanghai AI Laboratory; Tsinghua University; Fudan University(上海人工智能实验室; 清华大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出MARCH架构,通过内容路由状态锚点扩展循环记忆,在长序列上保持计算效率,经预训练后在多个长程任务中优于线性注意力变体,增强了循环长程记忆能力。

AI 中文摘要

Transformer的强大长上下文检索能力很大程度上归功于随上下文长度增长的token级记忆,但这种灵活性会在训练时带来二次计算复杂度,且在自回归推理时会产生线性增长的键值缓存。循环替代方案通过将全部历史压缩为固定大小的状态来提供高效解码,但在依赖召回的任务上表现通常不佳,因为早期关联常被后续更新覆盖,仅保留最近的上下文信息。本文提出Memory-Anchor Routing across Context History(MARCH),一种能将状态空间模型扩展至固定维度之外,同时在长序列上保持计算效率的网络架构。MARCH会周期性缓存累积的循环状态检查点作为状态锚点,并为每个锚点关联一个紧凑的、由内容调节的锚点键,这使MARCH能维护一个随上下文长度增长的记忆库,在历史分辨率与内存成本间提供可控权衡。在每个token处,MARCH生成一个锚点查询,以关注所有因果可用的状态锚点,输出通过沿当前状态对所有历史锚点进行注意力式聚合计算得到。我们表明,经标准预训练后,MARCH在常识推理、LongBench及上下文检索任务中,始终优于多种线性注意力变体。这些结果证明,内容路由状态缓存可大幅增强循环长程记忆,同时保留其原生计算路径。

英文摘要

Transformers owe much of their strong long-context retrieval capability to a token-level memory that grows with context length. This flexibility, however, incurs a quadratic computation complexity during training and a key--value cache that grows linearly during autoregressive inference. Recurrent alternatives offer efficient decoding by compressing the entire history into a fixed-size state, but often underperform on recall-intensive tasks since earlier associations usually get overwritten by subsequent updates, and only the most recent contextual information is retained. In this paper, we introduce Memory-Anchor Routing across Context History (MARCH), a network architecture that effectively scales state-space models beyond a fixed-size dimension, while maintaining computational efficiency over long-sequences. MARCH periodically caches cumulative recurrent-state checkpoints as state anchors and associates each anchor with a compact, content-conditioned anchor key. This lets MARCH maintain a memory bank, which can grow as context length increases, providing a controllable trade-off between historical resolution and memory cost. At each token, MARCH produces an anchor query to attend all causally available state anchors, and the output is calculated as an attention-style aggregation over all historical anchors along the current state. We show that after standard pretraining, MARCH consistently outperforms multiple linear attention variants across commonsense reasoning, LongBench, and in-context retrieval. These results demonstrate that content-routed state caching substantially strengthens recurrent long-range memory while preserving its native computation path.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑