arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15810cs.CV

VC-Attention:面向低位注意力的值平滑与Softmax铸造

VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

Xingyang Li, Dongyun Zou, Shining Zhang, Jiacheng Chen, Haocheng Xi, Lvmin Zhang, Jun-Yan Zhu, Song Han, Zhekai Zhang, Yujun Lin, Muyang Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对扩散Transformer低位注意力中值离群值损害精度和softmax高精度指数拖慢速度的问题,提出VC-Attention,通过值平滑与融合概率铸造,无需训练即提升保真度并加速1.46-3.6倍。

中文摘要 AI 辅助

扩散Transformer在视频生成领域达到了最先进的水平,但其长时空序列使得注意力成为部署成本的主要来源,因此一个可部署的低位内核必须既准确又快速。准确性受到离群值的限制:一个块的量化尺度由其最大条目决定,导致典型条目被限制在可表示值的狭窄范围内。先前的工作对查询和键进行平滑处理,但值离群值不遵循固定的通道或时空结构,仍然是输出误差的主要来源。速度受到softmax的限制:低位张量核心仅加速两次矩阵乘法,因此它们之间的高精度指数运算成为数据中心GPU上最长的流水线阶段。我们提出VC-Attention,一个无需训练的低位注意力框架,通过将值平滑与融合概率铸造配对来解决这两个问题。V-Smooth通过轻量级在线聚类重新排序值令牌,使得硬件块中的令牌能够良好地一起量化。它仅量化减去块均值后的残差,并从在线softmax已维护的行和恢复该均值。ExpCast-FP8通过一次融合乘加将对数域分数直接映射到E4M3概率代码,消除了FP32指数运算和格式转换。我们在B200、B300、H200、RTX PRO 6000和RTX 5090上实现了VC-Attention。在Wan2.2、LongCat-Video、HunyuanVideo-1.5和MiniMax-H3上,VC-Attention相比低位基线提高了保真度,在数据中心Blackwell和Hopper上将注意力内核速度比BF16 FlashAttention-4提升了1.46-1.59倍,在工作站显卡上提升了2.3-3.6倍,并且端到端生成片段的速度分别提升了1.13-1.19倍和1.36-1.70倍。

英文摘要

Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost, and a deployable low-bit kernel must be accurate and fast. Accuracy is limited by outliers: a block's quantization scale is set by its largest entries, leaving typical entries confined to a narrow range of representable values. Prior work smooths queries and keys, but value outliers follow no fixed channel or spatiotemporal structure and remain the dominant source of output error. Speed is limited by softmax: low-bit Tensor Cores accelerate only the two matrix multiplications, so the high-precision exponential between them becomes the longest pipeline stage on datacenter GPUs. We propose VC-Attention, a training-free low-bit attention framework that addresses both by pairing Value smoothing with a fused probability Cast. V-Smooth reorders value tokens by lightweight online clustering, so the tokens in a hardware block quantize well together. It quantizes only the residual after subtracting the block mean, and restores that mean from the row sum the online softmax already maintains. ExpCast-FP8 maps log-domain scores directly to E4M3 probability codes with one fused multiply-add, eliminating the FP32 exponential and the format conversion. We implement VC-Attention for B200, B300, H200, RTX PRO 6000, and RTX 5090. Across Wan2.2, LongCat-Video, HunyuanVideo-1.5, and MiniMax-H3, VC-Attention improves fidelity over low-bit baselines, speeds up the attention kernel over BF16 FlashAttention-4 by 1.46-1.59x on datacenter Blackwell and Hopper and by 2.3-3.6x on workstation cards, and generates a clip 1.13-1.19x and 1.36-1.70x faster end to end.

发表机构

  • Nunchux AI
  • UC Berkeley(加州大学伯克利分校)
  • Stanford(斯坦福大学)
  • CMU(卡内基梅隆大学)
  • MIT(麻省理工学院)
  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

相关深度报道

↑