预训练大语言模型中Softmax的近似:模型敏感性与核加速
Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration
浏览论文内容
中文总结 AI 辅助
针对预训练LLM中softmax近似,提出Rowmax-PoT及其硬件特化Rowmax-H15,在B200上实现FP8注意力前向加速12.4%-25.8%,能耗降低8.4%,困惑度影响极小。
中文摘要 AI 辅助
在NVIDIA Blackwell B200上,张量核心的吞吐量比特殊函数指数吞吐量高出两个数量级以上,这暴露了融合注意力核中指数评估的瓶颈。然而,预训练的Transformer可能不需要在每个元素上都精确评估指数。我们通过在十个冻结的仅解码器模型(0.5B-72B)中近似softmax来表征预训练模型的需求。softmax映射分配概率的位置数量以及行内分辨率可以大幅削减,但对相同位置进行均匀加权是有害的。固定分辨率预算的放置位置与其大小同样重要,且靠近行最大值处的分辨率始终更受青睐。在标量失真上匹配的扰动会产生符号相反的模型依赖性响应。这些发现促成了Rowmax-PoT,一种以每行最大值为锚点的粗粒度对数权重表示,以及Rowmax-H15,其在FlashAttention-4中的硬件特化。在B200上,修补后的FP8注意力前向传播在因果8K下主机端调用延迟测量中快12.4%,在非因果8K下快25.8%;在因果16K下,每次前向传播的板级能耗下降8.4%。在BF16核路径上单独测量(2K),Rowmax-H15在来自三个家族的五个模型上将困惑度提高了0.091-0.492%。
英文摘要
On NVIDIA Blackwell B200, tensor-core throughput outpaces special-function exponential throughput by more than two orders of magnitude, exposing exponential evaluation in fused attention kernels. A pretrained Transformer, however, may not need it evaluated accurately at every element. We characterize what a pretrained model does need by approximating softmax at inference in ten frozen decoder-only models (0.5B-72B). The number of positions the softmax map assigns probability to and within-row resolution can be cut substantially, yet uniform weighting of the same positions is damaging. Where a fixed resolution budget is placed matters as much as its size, with resolution near the row maximum consistently favored. Perturbations matched on scalar distortion produce model-dependent responses of opposite sign. These findings motivate Rowmax-PoT, a coarse logarithmic weight representation anchored at each row maximum, and Rowmax-H15, its hardware specialization in FlashAttention-4. On B200, the patched FP8 attention forward is 12.4% faster at causal 8K and 25.8% faster at non-causal 8K in host-side call-latency measurements; board energy per forward falls by 8.4% at causal 16K. Measured separately on the BF16 kernel path at 2K, Rowmax-H15 increases perplexity by 0.091-0.492% across five models from three families.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。