arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33051cs.LGcs.DC

SketchSSM: 写入完整状态,从紧凑草图中读取

SketchSSM: Write to the Full State, Read from a Compact Sketch

Omin Kwon, JoongWon Shin, Minseo Kim, Kurt Keutzer, Sehoon Kim, Jae W. Lee

AI总结:

SketchSSM通过离线固定基向量预计算状态读取输出并存储为紧凑草图,在解码时用查询系数重建结果,减少约10倍状态访问流量,同时保持准确率,并在NVIDIA B300上实现最高7.78倍内核加速。

AI中文摘要:

混合注意力模型用线性注意力取代了大部分softmax注意力层,减少了KV缓存增长,并使得更大的解码批次成为可能,其中循环状态访问成为主要瓶颈。ReplaySSM通过缓冲键和值来摊销状态更新,但每个新查询仍然需要完整状态读取,即使状态在状态更新之间保持不变。我们观察到,低秩状态加权查询近似能够准确保留状态读取输出。尽管未来查询未知,但用于近似它们的基向量可以离线固定。基于这一观察,我们引入了SketchSSM,它在保留完整状态更新的同时近似读取。在每次状态更新时,SketchSSM读取一次完整状态,为这些基向量预计算输出,并将其存储在紧凑草图中。每个后续解码步骤将草图向量与查询相关系数结合,以重建输出,而无需完整状态读取。在四个基于Mamba-2、GDN和KDA的模型中,SketchSSM将状态访问流量减少了约10倍,同时在四个解码基准测试中基本保持了平均准确率,并在四个RULER检索任务上保持了召回率。在一张NVIDIA B300上,与标准vLLM基线相比,线性注意力内核加速分别达到Mamba-2的7.78倍、GDN的5.22倍和KDA的5.20倍,在Nemotron 3 Super上的解码吞吐量最高提升2.64倍。

英文摘要:

Hybrid-attention models replace most softmax attention layers with linear attention, reducing KV-cache growth and enabling larger decode batches where recurrent-state access becomes a major bottleneck. ReplaySSM amortizes state updates by buffering keys and values, but each new query still requires a full-state read even though the state remains unchanged between state updates. We observe that low-rank state-weighted query approximation accurately preserves state-read outputs. Although future queries are unknown, the basis vectors used to approximate them can be fixed offline. Based on this observation, we introduce SketchSSM, which preserves full-state updates while approximating reads. At each state update, SketchSSM reads the full state once to precompute outputs for these basis vectors, storing them in a compact sketch. Each subsequent decode step combines the sketch vectors with query-dependent coefficients to reconstruct the output without a full-state read. Across four Mamba-2-, GDN-, and KDA-based models, SketchSSM at a mean sketch rank of 8 reduces state-access traffic by approximately 10$\times$ while matching the average accuracy of the FP32 full-state baseline across four decode benchmarks, and preserves recall on four RULER retrieval tasks. At this rank on one NVIDIA B300, linear-attention kernel speedups over the Standard vLLM baseline reach 7.30$\times$, 5.02$\times$, and 5.24$\times$ for Mamba-2, GDN, and KDA, respectively, with up to 2.77$\times$ higher decode throughput on Nemotron 3 Super.

↑