arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15533cs.DCcs.LG

DeltaLog:用于线性注意力解码的循环状态延迟物化

DeltaLog: Deferred Materialization of Recurrent States for Linear Attention Decoding

Junqing Lin, Jingwei Sun, Guangzhong Sun

首次发表
浏览论文内容

中文总结 AI 辅助

DeltaLog 是一种循环状态解码方案,通过延迟物化循环状态降低内存开销,在 GDN、KDA、RWKV6 上实现了更新内核加速、写流量减少及端到端服务速度提升。

中文摘要 AI 辅助

线性注意力模型通过将成对的 token 交互替换为循环状态更新,消除了 softmax 注意力的二次前缀计算和随上下文增长的 KV 缓存。然而,现有的解码实现通常会在每个生成的 token 之后物化并写回完整的循环状态,这使得状态维护成为内存流量的主要来源,尤其是对于具有大状态和多头的模型。本文提出了 DeltaLog,一种循环状态解码方案,在不改变模型语义的情况下降低了这种开销。具体而言,DeltaLog 将循环状态表示为一个密集基础状态,加上一个有限的近期紧凑更新日志。大多数解码步骤仅将紧凑更新因子追加到该日志中,而周期性合并步骤则将累积的更新折叠回密集基础状态。因此,模型观察到的密集状态与 eager 解码中的相同,但大多数完整状态写回被轻量级追加操作取代。我们为 GDN、KDA 和 RWKV6 实现了 DeltaLog,并将其集成到一个原型服务栈中。在这些模型上,DeltaLog 将循环状态更新内核加速了最高 1.86 倍,减少了已分析的循环状态写流量最高 7.83 倍,并相对于密集循环基线实现了 1.05 至 1.20 倍的端到端服务加速。

英文摘要

Linear attention models eliminate the quadratic prefix computation and context-growing KV cache of softmax attention by replacing pairwise token interactions with recurrent state updates. However, existing decoding implementations often materialize and write back the full recurrent state after every generated token, making state maintenance a major source of memory traffic, especially for models with large states and many heads. This paper presents DeltaLog, a recurrent-state decoding scheme that reduces this overhead without changing the model semantics. Specifically, DeltaLog represents the recurrent state as a dense base state together with a bounded log of recent compact updates. Most decode steps append only compact update factors to this log, while periodic merge steps fold the accumulated updates back into the dense base state. Thus, the model observes the same dense state as in eager decoding, but most full-state write-backs are replaced by lightweight append operations. We implement DeltaLog for GDN, KDA, and RWKV6 and integrate it into a prototype serving stack. Across these models, DeltaLog accelerates the recurrent-state update kernel by up to $1.86\times$, reduces profiled recurrent-state write traffic by up to $7.83\times$, and achieves $1.05$--$1.20\times$ end-to-end serving speedups over dense recurrent baselines.

发表机构

  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑