arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32712cs.AIcs.CL

MassAlloc 注意力:让注意力自行分配其计算量

MassAlloc Attention: Let Attention Allocate Its Own Compute

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

AI总结:

提出 MALA 融合注意力原语,根据归一化贡献自适应分配分数后计算,在保持 FullAttn 能力的同时显著降低计算量和延迟。

AI中文摘要:

FullAttn 通常会将极小的归一化质量分配给因果分数空间中的大部分区域,然而密集内核在形成每个 QK 分块后仍会执行完整的分数后路径。我们引入 MALA,一种融合注意力原语,它保留了对每个合法因果交互的分数访问,并利用归一化贡献来分配分数后计算。前向过程使用其动态更新的在线 softmax 归一化器,而反向过程则复用已确定的归一化器,仅利用标准注意力状态来推导嵌套的保留支持集。一个共同的容差控制训练和推理过程,允许对计算工作进行自适应保留。MALA 减少了低贡献的分数后计算。一项在 8K 长度下进行的工作量匹配研究隔离了分布自适应分配带来的收益:在总分数后工作量完全匹配的情况下,MALA 接近每个实例的参考质量预言机,平均省略质量为 0.0188%,而参考为 0.0182%。在从 1K 到 32K token 的上下文长度范围内,相同的容差相对于参考保持了较低的输出和梯度误差。在更广泛的控制性联想回忆对比中,随着上下文增长,MALA 紧密跟随 FullAttn,在 8K 时达到 89.67% 的准确率,而 FullAttn 为 89.97%。在 128K token 且使用张量并行的注意力算子基准测试中,相对于 FullAttn,MALA 在训练期间将前向和反向延迟分别降低了 2.2 倍和 3.0 倍,在推理期间将解码延迟降低了 1.6 倍。在从 0.6B 到 14B 参数的缩放定律训练中,MALA 在困惑度上紧密跟随 FullAttn,同时减少了总训练 FLOPs。由此产生的 14B 模型和 32B 模型(来自分别的持续训练)在知识、推理和长上下文检索得分上与 FullAttn 相当。这些结果表明,根据归一化注意力贡献来分配分数后计算,可以在减少注意力计算的同时保留 FullAttn 的已评估能力。

英文摘要:

FullAttn often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MALA, a fused attention primitive that preserves score access to every legal causal interaction and uses normalized contribution to allocate post-score computation. Forward uses its evolving online-softmax normalizer, while backward reuses the finalized normalizer to derive nested retained support using only standard attention state. A common tolerance governs training and inference, allowing for adaptive retention of the work. MALA reduces low-contribution post-score computation. A matched-work study at 8K isolates the benefit of distribution-adaptive allocation: under exactly matched total post-score work, MALA approaches a per-instance reference-mass oracle, with mean omitted mass of 0.0188% versus 0.0182%. Across context lengths from 1K to 32K tokens, the same tolerance maintains low output and gradient errors relative to the reference. Across a broader controlled associative-recall comparison, MALA closely tracks FullAttn as context grows, reaching 89.67% accuracy at 8K compared with 89.97% for FullAttn. In an attention-operator benchmark at 128K tokens with tensor parallelism, MALA reduces forward and backward latency during training by 2.2x and 3.0x and decoding latency during inference by 1.6x relative to FullAttn. Across scaling-law training from 0.6B to 14B parameters, MALA closely tracks FullAttn in perplexity while reducing total training FLOPs. The resulting 14B models and 32B models from separate continued training achieve comparable knowledge, reasoning, and long-context retrieval scores to FullAttn. These results indicate that allocating post-score computation according to normalized attention contributions can retain the evaluated capabilities of FullAttn while reducing attention computation.

↑