无损投机解码到底有多无损?Orthrus中数值精度所扮演的角色
Evaluating Losslessness in Speculative Decoding Under Finite-Precision Inference
AI总结:
本文复现Orthrus混合架构,发现其无损投机解码的精确轨迹匹配率在BF16下仅约45%,而在FP32下可达100%,表明无损性依赖于数值精度,且应与下游性能分开评估。
AI中文摘要:
Orthrus是一种混合自回归-扩散架构,通过使用冻结的自回归骨干网络并行生成多个token来加速自回归语言模型推理。其核心主张是,一种模型内共识机制能够实现无损投机解码,产生与自回归模型相同的输出序列。我们独立复现了Orthrus,并在不同数值精度下检验了这一主张。在BF16推理下,对于作者的检查点,在来自12个领域的1,190个提示中,仅有45%的情况实现了精确轨迹匹配;对于我们独立训练的模型,这一比例为43%。精确匹配的概率还与参考模型的响应条件困惑度密切相关。尽管存在这种轨迹差异,Orthrus在下游lm-eval-harness基准上并未表现出系统性退化。相比之下,使用FP32重复轨迹评估时,所有评估提示均实现了精确轨迹匹配。这些结果表明,Orthrus的实际无损性取决于数值精度,且精确轨迹等价性应与下游任务性能分开评估。
英文摘要:
Lossless speculative decoding is typically defined at the algorithmic level: a speculative procedure proposes multiple tokens and a verification procedure is designed to preserve the output trajectory of an autoregressive reference model exactly. In practical neural inference, however, this guarantee is implemented using finite-precision floating-point computations, and discrete token selection can amplify small numerical differences into divergent generation trajectories. We investigate this distinction using Orthrus, a hybrid autoregressive-diffusion architecture that performs self-drafting and self-verification within a frozen autoregressive backbone, as a representative case study. Across 1,190 prompts from 12 domains, exact trajectory matching under BF16 occurs for only 45\% of the authors' checkpoint generations and 43% of those from our independently trained model. The probability of matching is strongly associated with the response-conditional perplexity of the autoregressive reference, indicating that trajectory divergence is not uniform across inputs. Despite these divergences, Orthrus does not exhibit systematic degradation on the evaluated downstream tasks. In contrast, FP32 inference yields exact trajectory matching on all evaluated prompts. These results demonstrate a gap between algorithmic losslessness and its implementation under finite-precision arithmetic, and motivate evaluating lossless speculative decoding at the level of exact generation trajectories as well as downstream task performance.