arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26621cs.LGcs.AI

贪心解码并非精度不变:LLM 推理中的跨精度输出分歧

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

  • University of Tennessee, Knoxville(田纳西大学诺克斯维尔分校)
  • University of Chicago(芝加哥大学)
  • Amazon(亚马逊)
  • University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li

AI总结:

本研究揭示LLM贪心解码在BF16与FP16下输出分歧,提出选择性FP32 LM头部重计算干预,以低延迟开销显著提升跨精度一致性,但仅部分缓解。

AI中文摘要:

从大型语言模型进行贪心解码通常被视为确定性的。我们表明它并非精度不变的:在相同硬件上,同一模型、提示和解码算法在 BF16 与 FP16 下会产生不同的输出。在我们对六个模型(1.1B-7B 参数,四个系列;分歧在 12B 规模下另有表征)和三个基准的评估中,49-100% 的提示发生分歧;单个 token 翻转常常级联为轨迹级分歧。我们开发了一种经验性误差传播分析,发现 22 层累积的主体误差无法区分翻转与非翻转步骤;结果主要取决于 LM 头部的 top-2 logit 边际相对于 top-2 候选之间方向性扰动的比较。该分析对干预结果做出了五个可测试的预测,包括应用更多 FP32 计算(更广范围)会使一致性变差。实验与所有五个预测相符。我们评估的最佳低开销干预措施——仅在边际低于阈值时触发的选择性 FP32 LM 头部重计算——在低批量(批量大小≤4)单流推理中,以低于 4% 的延迟开销,在 A10G 上实现了 +22-36 个百分点的精确一致性提升(在 L4 和 A100 上为 +12-21 个百分点)。我们绘制了跨六个模型和四种批量大小的适用性边界,并假设训练时的精度稳定性是一个决定性因素。该方法是一种部分缓解措施,而非通用的确定性保证:当主体来源的误差占主导时,其益处消失,包括在我们的测试中批量大小≥8 以及端到端 FP8 的情况下。

英文摘要:

Greedy decoding from large language models is commonly treated as deterministic. We show it is not precision-invariant: the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware. Across our evaluations of six models (1.1B-7B parameters, four families; divergence additionally characterised at 12B) and three benchmarks, 49-100\% of prompts diverge; a single token flip often cascades into trajectory-level divergence. We develop an empirical error-propagation analysis and find that 22 layers of accumulated body error do not distinguish flipping from non-flipping steps; the outcome depends primarily on the top-two logit margin at the LM head relative to the directional perturbation between the top-two candidates. The analysis makes five testable predictions about intervention outcomes, including that applying more FP32 compute (broader scope) makes agreement worse. The experiments match all five predictions. The best-performing low-overhead intervention we evaluate, selective FP32 LM head recomputation, triggered only when the margin falls below a threshold, delivers +22-36 pp exact agreement on A10G (+12-21 pp on L4 and A100) at less than 4\% latency overhead in low-batch (batch size <=4) single-stream inference. We map the applicability boundary across six models and four batch sizes, and hypothesise that training-time precision stability is a determining factor. The method is a partial mitigation rather than a universal determinism guarantee: its benefit vanishes when body-originated error dominates, including at batch size >=8 and under end-to-end FP8 in our tests.

补充信息

↑