arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19758cs.CL

FlashPrefill V2:面向长上下文大语言模型服务的块稀疏预填充注意力机制

FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

Qihang Fan, Huaibo Huang, Zhiying Wu, Bingning Wang, Ran He

首次发表
浏览论文内容

中文总结 AI 辅助

FlashPrefill V2通过均值校正、适配FlashAttention-3/4的算子设计及原生支持分页KV缓存等,在128K上下文长度下为长上下文LLM服务提供了高效的块稀疏预填充注意力方案,加速比显著。

中文摘要 AI 辅助

长上下文建模是大语言模型的关键能力,但注意力机制的二次复杂度仍是关键瓶颈,尤其在计算密集型的预填充阶段。我们之前的工作FlashPrefill通过瞬时模式发现和基于最大值的动态阈值缓解了该成本,但它仍是距离生产部署较远的算法原型。本文提出FlashPrefill V2,从三个维度将FlashPrefill从原型推进至实用长上下文服务:第一,引入均值校正项有效抑制近似误差,即使在极端稀疏度下也能将性能下降控制在可接受范围;第二,通过PackGQA内存访问、warp专业化和乒乓流水线重新设计稀疏注意力算子,完全适配最新的FlashAttention-3/4实现,并支持FP8推理以满足实用量化需求;第三,FlashPrefill V2原生支持分页KV缓存和连续批处理,可作为注意力后端集成至SGLang等现代推理框架。在NVIDIA H20 GPU(部署最广泛的推理加速器之一)上的大量评估表明,在128K上下文长度下,FlashPrefill V2在FP8和BF16精度下分别实现了比FlashAttention-2高47.26倍和27.19倍的加速,在FP8精度下,相较于适配FA3/4的密集基线仍实现了30.49倍的加速。

英文摘要

Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. Our previous work, FlashPrefill, mitigates this cost through instantaneous pattern discovery and max-based dynamic thresholding; however, it remains an algorithmic prototype that is still distant from production deployment. In this paper, we present FlashPrefill V2, which evolves FlashPrefill from a prototype toward practical long-context serving along three dimensions. First, we introduce a mean correction term that effectively suppresses the approximation error, keeping performance degradation manageable even at extreme sparsity levels. Second, we redesign the sparse attention operator with PackGQA memory access, warp specialization, and pingpong pipelining, fully aligning with the latest FlashAttention-3/4 implementations and supporting FP8 inference to meet practical quantization requirements. Third, FlashPrefill V2 natively supports paged KV cache and continuous batching, allowing integration as an attention backend in modern inference frameworks such as SGLang. Extensive evaluations on NVIDIA H20 GPUs---among the most widely deployed inference accelerators---demonstrate that FlashPrefill V2 delivers up to 47.26x and 27.19x speedups over FlashAttention-2 at 128K context length under FP8 and BF16 precision, respectively, and, in FP8, still achieves a 30.49x speedup against an FA3/4-aligned dense baseline.

发表机构

  • MAIS&NLPR, CASIA(中国科学院自动化研究所多模态人工智能系统与国家重点实验室)
  • UCAS(中国科学院大学)
  • WeChat, Tencent(腾讯微信)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑