arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33135cs.DC

MpFA:Blackwell GPU 上硬件高效的免训练 QK4V8 FlashAttention 内核

MpFA: Hardware-Efficient Train-Free QK4V8 FlashAttention Kernels on Blackwell GPUs

  • College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机学院)

机构由 AI 辅助整理,请以论文原文为准。

Chencheng Deng, Jianbin Fang, Dezun Dong

AI总结:

针对长上下文 LLM 推理的注意力瓶颈,提出免训练混合精度 FlashAttention 内核 MpFA,采用 QK4PV8 与秩一补偿,在 B200 上显著提升吞吐并恢复精度。

AI中文摘要:

长上下文 LLM 推理将现代 GPU 服务栈推向注意力受限的状态,其中计算和内存均由 softmax-GEMM 流水线主导。在 NVIDIA Blackwell GPU 上,FP4 Tensor Core 提供高矩阵乘法吞吐量,但我们发现全 FP4 注意力往往无法将这种吞吐量转化为端到端的加速,原因在于非矩阵乘法成本:softmax 后的在线量化、张量/共享内存数据移动以及 softmax 路径上的争用。我们提出了 MpFA,一个针对 Blackwell 优化的免训练 FlashAttention 内核。在硬件特征分析的指导下,MpFA 使用混合精度:QK 使用 NVFP4,PV 使用 FP8(QK4PV8)。这保留了低比特 QK 的吞吐量,同时避免了 FP4 PV 的转换和缩放开销。为了在不进一步加重 softmax 流水线负担的情况下恢复精度,MpFA 引入了秩一平滑补偿,以额外的 Tensor Core MMA 实现。MpFA 还通过细粒度异步流水线、张量内存重用以及跨预填充和解码阶段的自适应并行分区进一步提高了性能。在 NVIDIA B200 上,跨越 16K-128K 上下文长度,MpFA 相对于最先进的 BF16/FP8 基线提高了预填充吞吐量,并在 Llama-3.1-8B 和 Qwen3-14B 上将端到端输出吞吐量相对于 BF16 FA4 提高了 2.81 倍。在五个基准套件和两个模型上,秩一补偿恢复了 62.5% 的精度损失,内核开销约为 2.0%。

英文摘要:

Long-context LLM inference pushes modern GPU serving stacks into an attention-bound regime, where both compute and memory are dominated by the softmax-GEMM pipeline. On NVIDIA Blackwell GPUs, FP4 Tensor Cores offer high matmul throughput, but we find that fully FP4 attention often fails to translate this throughput into end-to-end speedups due to non-matmul costs: online quantization after softmax, tensor/shared-memory data movement, and contention on the softmax path. We present MpFA, a training-free FlashAttention kernel optimized for Blackwell. Guided by hardware characterization, MpFA uses mixed precision: NVFP4 for QK and FP8 for PV (QK4PV8). This preserves low-bit QK throughput while avoiding the conversion and scaling overheads of FP4 PV. To recover accuracy without further stressing the softmax pipeline, MpFA introduces rank-one smoothing compensation implemented as an additional Tensor Core MMA. MpFA further improves performance with a fine-grained asynchronous pipeline, tensor-memory reuse, and adaptive parallel partitioning across prefill and decode. On an NVIDIA B200 and across 16K-128K contexts, MpFA improves prefill throughput over state-of-the-art BF16/FP8 baselines and increases end-to-end output throughput by 2.81$\times$ over BF16 FA4 across Llama-3.1-8B and Qwen3-14B. Across five benchmark suites and two models, rank-one compensation recovers 62.5% of the accuracy loss with about 2.0% kernel overhead.

↑