arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LampAttention:面向专用加速器的前瞻混合精度FlashAttention

LampAttention: Look-Ahead Mixed-Precision FlashAttention for Dedicated Accelerators

Stanislav Budzinskiy, Marian Gloser, Tolunay Yilmaz, Ying Hong Tham, Yuanyi Lin, Wenyi Fang, Fan Wu, Philipp Petersen

arXiv 2609.39361首次发表:更新:

发表机构

University of Vienna; Huawei Heisenberg Research Center; Huawei Technologies Co. Ltd(维也纳大学; 华为海森堡研究中心; 华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LampAttention,一种混合精度FlashAttention的硬件-算法协同设计,通过8位计算与自适应16位重算敏感子块,在专用加速器上高效恢复模型性能。

AI 中文摘要

虽然大多数注意力logits可以在低精度下计算而不会降低数值稳定性,但当前的注意力内核未能利用这一现象。我们提出了一种新颖的硬件-算法协同设计,即混合精度FlashAttention。我们的方法以8位格式累积键-查询乘积并评估其指数,然后自适应地识别敏感子块并以16位格式重新计算它们。我们提出了一个能够高效执行此流程的专用加速器的规格说明。使用Qwen3和Gemma 3进行的模拟实验表明,将选择性的少数子块重新路由到高精度足以恢复基线模型性能。

英文摘要

While most attention logits can be computed in low precision without degrading numerical stability, current attention kernels fail to exploit this phenomenon. We introduce a novel hardware-algorithm co-design in the form of mixed-precision FlashAttention. Our method accumulates key-query products and evaluates their exponentials in 8-bit formats, then adaptively identifies sensitive sub-blocks and recomputes them in 16-bit formats. We propose the specifications for a dedicated accelerator capable of executing this pipeline efficiently. Simulated experiments with Qwen3 and Gemma 3 show that rerouting a selective minority of sub-blocks to high precision is sufficient to recover the baseline model performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑