加速跨GPU架构的LLM推理非确定性缓解
Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures
查看机构详情
- Georgia Institute of Technology(佐治亚理工学院)
- University of California, Merced(加州大学默塞德分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究提出固定配置的融合升型GEMM内核,通过固定浮点归约顺序实现跨GPU架构的LLM推理逐位可复现,同时提升1.17-3.1倍速度并减半内存流量。
中文摘要 AI 辅助
大型语言模型(LLM)的输出在贪心解码下预期是可复现的,然而在实践中,相同的模型、提示词和软件栈在不同GPU上会产生不同的输出。根本原因是浮点非结合性与硬件相关的内核选择相结合。推理框架在每个架构上选择不同的矩阵乘法内核,具有不同的并行归约顺序和未指定的张量核心算术,由此产生的舍入差异可能翻转输出令牌。现有解决方案的跨架构可复现性不完善,并且会带来显著的性能损失。我们提出了一种解决方案,采用一组固定配置的融合升型GEMM内核,从内存加载16位权重,在寄存器中将其升型为FP32,并使用IEEE-754算术以归约顺序进行累加,该归约顺序是问题形状的纯函数,因此与设备、其SM数量或内核调度无关。通过将浮点归约顺序固定为仅取决于问题形状的函数,每个GPU都执行相同的操作序列,因此线性层的跨架构可复现性归结为正确的IEEE-754算术,而不是舍入差异保持在平局翻转阈值以下。我们确认我们的解决方案的线性层输出在NVIDIA Ampere、Ada和Hopper GPU上逐位相同,同时端到端运行速度比最先进的解决方案快1.17至3.1倍,并将权重内存流量减半。
英文摘要
Large language model (LLM) outputs are expected to be reproducible under greedy decoding, yet in practice the same model, prompt, and software stack produce different outputs on different GPUs. The root cause is floating-point non-associativity combined with hardware-dependent kernel selection. Inference frameworks select different matrix-multiplication kernels on each architecture, with different parallel reduction orders and unspecified tensor-core arithmetic, and the resulting rounding differences can flip output tokens. Existing solutions have imperfect cross-architecture reproducibility and incur a significant performance penalty. We present a solution employing a set of fixed-configuration fused-upcast GEMM kernels that load 16-bit weights from memory, upcast them to FP32 in registers, and accumulate with IEEE-754 arithmetic in a reduction order that is a pure function of the problem shape and is therefore independent of the device, its SM count, or kernel scheduling. By fixing the floating-point reduction order as a function of problem shape alone, every GPU runs the same operation sequence, so cross-architecture reproducibility of the linear layers reduces to correct IEEE-754 arithmetic rather than to rounding differences staying below a tie-flip threshold. We confirm our solution's linear-layer outputs are bitwise identical across NVIDIA Ampere, Ada, and Hopper GPUs, while running $1.17$ to $3.1\times$ faster end-to-end than the state-of-the-art solution and cutting weight-memory traffic in half.