arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02947cs.LGcs.CL

ATFlash:用于计算/内存高效LLM推理的逐RoPE波长注意力窗口

ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference

Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri

首次发表
浏览论文内容

中文总结 AI 辅助

ATFlash提出逐RoPE波长注意力窗口,修剪查询-键内积项,在保持长上下文任务性能的同时,可移植至FlashAttention-4等框架,在多款模型上实现LLM推理的计算与内存效率提升。

中文摘要 AI 辅助

带旋转位置嵌入(RoPE)的注意力分数可精确分解为二维旋转频率对的和,每对的波长决定其区分位置的最大距离。基于此结构,我们提出逐RoPE波长距离窗口:该窗口会修剪查询-键内积项中超过波长比例距离的部分。与滑动窗口不同,每个键至少可通过低频对被访问,且修剪率与输入无关,具有与序列长度N成对数关系的闭式表达式,这与MInference等动态稀疏方法形成对比。这种token级选择与我们的频率级修剪正交,因此该窗口可应用于这些方法之上。在Qwen2.5-0.5B和Llama-3.2-3B上,该窗口在模型原生上下文长度内修剪了37%-48%的查询-键内积项;相较于全注意力,在LongBench-v2上下文上,top-1匹配率保持在96%-98%,平均输出分布KL值处于10^-3 nat级别。我们在RULER、OpenAI-MRCR、LongCodeQA和∞Bench等长上下文基准上测试绝对分数,发现其基本保持。我们将该窗口实现为查询-键收缩轴的切片,保留在线-softmax循环不变,并以最小差异移植到已发布的FlashAttention-4预填充和FlashInfer解码中。在搭载Llama的RTX PRO 6000上,两个移植版本均优于原版,且增益随上下文长度增长,在128K时达到1.29倍加速;在Qwen2.5-7B-1M上,修剪57%的内积项后,1M-token上下文下的端到端加速达1.31倍。

英文摘要

The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelength limits how far it can discriminate position. Aligned with this structure, we propose the per-RoPE-wavelength distance window: it prunes the query--key inner-product terms beyond a wavelength-proportional distance. Unlike a sliding window, every key remains reachable, at least through the low-frequency pairs. The reduction rate is input-independent, with a closed form logarithmic in the sequence length $N$, in contrast to dynamic-sparse methods like MInference. Such token-level selection is orthogonal to our frequency-level pruning. The window can therefore be applied on top of those methods. On Qwen2.5-0.5B and Llama-3.2-3B, the window prunes 37--48\% of the query--key inner-product terms within each model's native context length. Relative to full attention, the top-1 match rate stays at 96--98\% and the mean output-distribution KL at the $10^{-3}$-nat level on LongBench-v2 contexts. We examine absolute scores on long-context benchmarks such as RULER, OpenAI-MRCR, LongCodeQA, and $\infty$Bench: they are broadly preserved. We implement the window as a slice of the query--key contraction axis, leaving the online-softmax recurrences untouched, and port it with minimal diffs into the released FlashAttention-4 prefill and FlashInfer decode. On RTX PRO 6000 with Llama, both ports outpace stock with gains growing with context length, up to $1.29\times$ at 128K. End to end on Qwen2.5-7B-1M, with 57\% of the inner-product terms pruned, the speedup reaches $1.31\times$ at a 1M-token context.

↑