arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00049cs.LG

面向大语言模型的快速多项式超越函数

Fast Polynomial Transcendentals for LLMs

Robert Hu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对GPU硬件演进中特殊函数单元瓶颈,提出用短多项式程序替代LLM中的sigmoid、tanh等超越函数,在GB200上实现最高8%的训练吞吐提升,且模型行为影响极小。

中文摘要 AI 辅助

图形处理单元(GPU)的各代产品以不同速率扩展矩阵、特殊函数和内存流水线,因此内核瓶颈会随着硬件演进而变化。FlashAttention-4 在 NVIDIA Blackwell 上暴露了注意力机制内部的这种不平衡。我们测试了短多项式程序能否加速大语言模型(LLM)中的其他特殊函数单元(SFU)操作。我们首先在隔离的 IEEE 二进制16(FP16)扫描中,将原生 PyTorch 评估与打包的融合乘加(FMA)程序进行比较,该扫描覆盖了 L2 常驻和高带宽内存(HBM)常驻的工作集。然后,我们在四个 GB200 集成任务中,用三次或四次 bfloat16(BF16)程序替换原生 sigmoid、tanh 和 sigmoid 线性单元(SiLU):密集 SiLU、tanh 软上限注意力、sigmoid 注意力和路由专家 Swish 门控线性单元(SwiGLU)。这些程序在消费内核内结合了解析对称性、目标格式舍入和打包算术。隔离路径在 L2 中提升了 1.19--2.19 倍,在 HBM 中提升了 1.00--1.70 倍。密集 SiLU、tanh 软上限注意力和路由专家替换分别将完整训练步骤的吞吐量提升了 2.7%、2.9% 和 8.0%。sigmoid 注意力替换将完整注意力前向传播提升了 7.4%,完整 GPU 步骤提升了 0.3%。同检查点开放权重消融和每任务一次配对预训练比较将评估扩展到模型行为。在接近 1000 亿 token 的常见视野下,四个任务的最终平滑训练损失差异(多项式减原生)范围从 $-0.107$ 到 $+0.079$。

英文摘要

Graphics processing unit (GPU) generations scale matrix, special-function, and memory pipelines at different rates, so kernel bottlenecks move as hardware evolves. FlashAttention-4 exposed this imbalance inside attention on NVIDIA Blackwell. We test whether short polynomial programs can accelerate other special-function-unit (SFU) operations in large language models (LLMs). We first compare native PyTorch evaluation with packed fused multiply--add (FMA) programs in an isolated IEEE binary16 (FP16) sweep spanning L2-resident and high-bandwidth-memory (HBM)-resident working sets. We then replace native sigmoid, tanh, and sigmoid linear unit (SiLU) with degree-3 or degree-4 bfloat16 (BF16) programs in four GB200 integration tasks: dense SiLU, tanh-softcapped attention, sigmoid attention, and routed-expert Swish-gated linear unit (SwiGLU). The programs combine analytical symmetry, target-format rounding, and packed arithmetic inside consuming kernels. The isolated paths improve by 1.19--2.19x in L2 and 1.00--1.70x in HBM. The dense-SiLU, tanh-softcapped-attention, and routed-expert substitutions improve complete training-step throughput by 2.7\%, 2.9\%, and 8.0\%, respectively. The sigmoid-attention substitution improves complete-attention forward by 7.4\% and the complete GPU step by 0.3\%. Same-checkpoint open-weight ablations and one paired pre-training comparison per task extend the evaluation to model behavior. At common horizons near 100 billion tokens, the final smoothed training-loss differences (polynomial minus native) range from $-0.107$ to $+0.079$ across the four tasks.

↑