arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

低精度Transformer推理中的随机舍入:小型GPT-2的可变精度模拟研究

Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2

Yohan Chatelain, Pablo de Oliveira Castro

arXiv 2610.01889首次发表:更新:

发表机构

Krembil Centre for Neuroinformatics; Centre for Addiction and Mental Health; Université Paris-Saclay; UVSQ(克雷姆比尔神经信息学中心; 成瘾与心理健康中心; 巴黎萨克雷大学; 凡尔赛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过可变精度模拟,发现低精度Transformer推理中随机舍入与就近舍入的效果取决于网络位置,混合配置可显著降低困惑度。

AI 中文摘要

低精度Transformer推理应使用随机舍入(SR)还是就近舍入(RN)?答案取决于在网络中的位置。我们通过固定数值格式、仅改变各个操作位点的舍入规则来隔离这一效应。为了在任意选择的精度下进行实验,我们通过可变精度随机舍入(VPSR)算法将PRISM向量化舍入库扩展到任意虚拟精度,并证明舍入决策在硬件浮点中精确评估。我们开发了两种分析,为这种位点级权衡提供互补的见解。首先,线性投影的概率前向误差界表明,SR的误差包络随归约长度n按O(√n u)增长,而RN为O(n u),这一差距在低精度下迅速扩大,在长多层感知器(MLP)下投影中最为显著。其次,输出softmax处期望交叉熵损失变化的二阶分解,分为有符号漂移、漂移曲率和Fisher加权方差惩罚,揭示了两个位点行为相反的原因:MLP噪声主要是均匀的logit偏移,softmax对其不变,因此SR的方差在很大程度上被折扣;头部噪声在词汇表上非均匀,因此不被折扣。在DistilGPT-2的t=6有效位下,观测结果与理论一致:MLP中的SR将困惑度提高到全精度参考值的1.15倍,而RN为2.21倍。在语言模型头部,顺序反转,因为SR引入非均匀方差,而确定性RN不引入。在混合精度配置(MLP输出t=6)中,将SR分配给MLP、RN分配给头部,困惑度达到全精度参考值的1.10倍以内,比匹配位数的RN降低了28%。

英文摘要

Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments at freely chosen precisions, we extend the PRISM vectorized rounding library to arbitrary virtual precision via a variable-precision stochastic rounding (VPSR) algorithm, proving that the rounding decision is evaluated exactly in hardware floating point. We develop two analyses providing complementary insight into this site-level trade-off. First, a probabilistic forward-error bound for linear projections shows that SR's error envelope grows as $O(\sqrt{n} u)$ in reduction length $n$, versus $O(n u)$ for RN, a gap that widens rapidly at low precision and is most pronounced in the long multilayer perceptron (MLP) down-projection. Second, a second-order decomposition of expected cross-entropy loss change at the output softmax into signed drift, drift curvature, and a Fisher-weighted variance penalty reveals why the two sites behave oppositely: MLP noise is predominantly a uniform logit shift to which softmax is invariant, so SR's variance is largely discounted; head noise is non-uniform across the vocabulary and is not. On DistilGPT-2 at $t=6$ significand bits, observations match theory: SR in the MLP raises perplexity to 1.15x the full-precision reference, versus 2.21x for RN. At the language-model head, the ordering reverses because SR introduces non-uniform variance, whereas deterministic RN carries none. In a mixed-precision configuration (MLP output at $t=6$), assigning SR to the MLP and RN to the head brings perplexity within 1.10x of the full-precision reference, a 28% reduction over matched-bit RN.

Comments35 pages, 10 figures, 4 tables. Code and evaluation pipeline available at https://github.com/big-data-lab-team/fuzzy-llm and archived on Zenodo at https://doi.org/10.5281/zenodo.23066028

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑