arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MOIRA:面向长上下文解码的基于质量的带不规则注意力的索引

MOIRA: Mass-Oriented Indexing with Ragged Attention for Long-Context Decoding

Dich Nhat Minh Nguyen, Tran Dang Duong Nguyen

arXiv 2610.04313首次发表:更新:

AI 中文总结

MOIRA提出一种无需训练的稀疏解码方法,通过按KV头和层自适应页预算及自规划注意力内核,在保持精度的同时大幅降低长上下文解码延迟并提升吞吐量。

AI 中文摘要

长上下文解码受限于内存带宽,因为每个输出token都会读取每一层的KV缓存。稀疏解码通过仅读取部分KV缓存来降低这一成本。我们观察到,查询所需的页数在不同KV头、层和步骤之间差异很大。固定预算虽然简单,但它是针对高需求情况设计的,并且需要针对每个工作负载进行调整;自适应预算能更灵活地跟随这种变化,但现有设计为此付出了额外的选择成本或训练成本。在内核层面,FlashAttention-3(FA3)和FlashInfer是为长度相似的行设计的:当页列表长度因KV头而异时,它们要么填充列表(牺牲大部分稀疏节省),要么导致线程块负载不均衡,要么依赖在CUDA图之外运行的主机端计划。我们提出了MOIRA,一种在vLLM中无需训练的稀疏解码路径,其预算针对每个KV头和每层自适应。对于每个请求、层、KV头和步骤,一个覆盖规则保留最小的页集合,其估计注意力质量达到分数γ。一个新的内核,自规划注意力,让每个线程块根据列表长度自行推导其工作份额,从而使整个解码步骤保持在CUDA图内。在H200上,在RULER的128k上下文下,γ=0.99的MOIRA在读取约30%页面的同时匹配稠密精度,并将每个输出token的时间(TPOT)相对于稠密FA3降低了2.2-2.5倍;γ=0.98时,TPOT降低2.7倍,且保持在稠密噪声范围内。在高服务负载下,吞吐量提升高达51%。这些结果表明,针对每个头和层自适应的预算,配合将此类预算保持在CUDA图内的内核,使稀疏解码既灵活又快速。

英文摘要

Long-context decoding is limited by memory bandwidth, because every output token reads the KV cache of every layer. Sparse decoding reduces this cost by reading only part of the KV cache. We observe that the number of pages a query needs varies widely across KV heads, layers and steps. Fixed budgets are simple, but they are sized for demanding cases and tuned per workload; adaptive budgets follow this variation more flexibly, but existing designs pay for it with extra selection cost or training. At the kernel level, FlashAttention-3 (FA3) and FlashInfer are designed for rows of similar length: with page lists whose length differs per KV head, they either pad the lists (forfeiting much of the sparse saving), leave thread blocks unbalanced, or rely on a host-side plan that runs outside the CUDA graph. We propose MOIRA, a training-free sparse decode path in vLLM whose budget adapts per KV head and per layer. For every request, layer, KV head and step, a coverage rule keeps the smallest set of pages whose estimated attention mass reaches a fraction $γ$. A new kernel, self-planning attention, lets each thread block derive its own share of the work from the list lengths, so the whole decode step stays inside the CUDA graph. On an H200, at RULER's 128k context, MOIRA with $γ=0.99$ matches dense accuracy while reading about 30% of the pages and reduces the time per output token (TPOT) by 2.2-2.5$\times$ relative to dense FA3; with $γ=0.98$ it reduces TPOT by 2.7$\times$ and stays within the noise of dense. Under high serving load it raises throughput by up to 51%. These results suggest that a budget adapted per head and layer, paired with a kernel that keeps such budgets inside the CUDA graph, makes sparse decoding both flexible and fast.

Comments19 pages, 7 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑