arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

测量与减少语言模型中的跨厂商不匹配

Measuring and Reducing Cross-Vendor Mismatch in Language Models

Erland Hilman Fuadi, Chong Tian, Xiaosong Ma, Qirong Ho

arXiv 2610.05458首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究测量并减少语言模型在GPU厂商间的输出不匹配,发现累加顺序是根源,提出部分精度提升方案,并指出单一度量评估的局限性。

AI 中文摘要

在不同图形处理单元(GPU)厂商上运行相同的语言模型,即使模型权重和输入相同,也可能产生不同的逻辑值(logits)。我们使用五个度量族分析了两个密集模型和两个混合专家(MoE)模型中的跨厂商不匹配,即位级相等性、逻辑值差异、Top-K一致性、令牌一致性和任务准确性。我们将不匹配的一个来源追溯到厂商矩阵指令内部的累加顺序。将精度提升至FP32可将密集模型的逻辑值误差降低43%,但运行时间增加三倍;然而,仅将MLP层保持在BF16精度,就能在运行时间增加1.3倍的情况下保留此改进的94%,因此完全提升至FP32的大部分成本收益甚微。在MoE模型中,FP32和FP16均降低了概率误差,但增加了逻辑值误差并改变了专家选择,且FP16在密集模型中失效。输出头低秩适配器(LoRA)也无济于事,因为最终隐藏状态无法预测不匹配。这种不匹配还会延续到训练中。在固定所有随机种子的情况下,从在AMD上运行的教师模型蒸馏出的学生模型,与从在NVIDIA上运行的同一教师模型蒸馏出的学生模型,在431个MMLU问题上的回答不同。在FP32提升精度下,位级相等性几乎不变,而输出分布大部分向参考分布移动,因此用单一度量判断跨厂商一致性会误读其成本和收益。代码可在以下网址获取:此https URL。

英文摘要

Running the same language model on different graphics processing unit (GPU) vendors can produce different logits, even when the model weights and inputs are the same. We analyze cross-vendor mismatch in two dense and two mixture-of-experts (MoE) models with five metric families, namely bitwise equality, logit differences, top-K consistency, token agreement, and task accuracy. We trace one source of the mismatch to accumulation order inside vendors' matrix instructions. Upcasting to FP32 reduces the dense model's logit error by 43% at three times the runtime, yet keeping only the MLPs in BF16 retains 94% of this gain at 1.3 times the runtime, so most of the cost of full upcasting buys little. In the MoE models, FP32 and FP16 both lower the probability error but raise the logit error and change expert selection, and FP16 fails in the dense model. An output-head low-rank adapter (LoRA) does not help either, since the final hidden state does not predict the mismatch. The mismatch also carries into training. With every seed fixed, a student distilled from a teacher running on AMD answers 431 MMLU questions differently from one distilled from the same teacher on NVIDIA. Under FP32 upcasting, bitwise equality barely changes while the output distributions move most of the way to the reference, so judging cross-vendor agreement by a single measure misreads both its cost and its gains. Code is available at https://github.com/crova-project/crova.

Comments23 pages, 9 figures, 15 tables. Code: https://github.com/crova-project/crova

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑