FastGuide:加速扩散大语言模型的奖励引导
FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models
浏览论文内容
中文总结 AI 辅助
FastGuide提出一种并行与自回归自适应混合解码方法,加速扩散语言模型的奖励引导,在三个基准上实现高达4.4倍加速且保持生成质量。
中文摘要 AI 辅助
基于梯度的奖励引导提供了一种灵活的方式,在推理时使用下游奖励模型来控制掩码扩散语言模型。然而,其计算成本仍然很高,因为每次解码迭代都需要昂贵的扩散模型前向传播和奖励模型反向传播步骤。为了解决这个问题,我们引入了FastGuide,一种并行和自回归解码的自适应混合方法,以加速扩散语言模型的奖励引导。与并行解码类似,FastGuide通过每个解码步骤计算一次引导并复用它来生成多个令牌,从而分摊奖励模型反向传播的成本。在每个解码步骤内,FastGuide通过一次取消掩码一个令牌来使扩散前向传播自回归化,同时利用KV缓存技术和注意力的稀疏重计算,在每次取消掩码后高效地重新计算令牌分布。最后,为了使混合解码适应模型的置信度,FastGuide会推迟任何模型在其重新计算的分布下不自信的令牌。在三个奖励基准上的实验表明,FastGuide比顺序奖励引导解码快高达4.4倍,同时保持相似的生成质量。
英文摘要
Gradient-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time. However, its computational cost remains high as each decoding iteration incurs expensive diffusion model forward passes and reward model backpropagation steps. To address this, we introduce FastGuide, an adaptive hybrid of parallel and autoregressive decoding to accelerate reward guidance for diffusion language models. In analogy to parallel decoding, FastGuide amortizes the cost of reward model backpropagation by computing guidance once per decoding step and reusing it to generate multiple tokens. Within each decoding step, FastGuide makes diffusion forward passes autoregressive by unmasking tokens one at a time while efficiently recomputing token distributions after each unmasking by utilizing KV caching techniques and sparse recomputation of attention. Lastly, to adapt hybrid decoding to the model's confidence, FastGuide defers any token that the model is unconfident about under its recomputed distribution. Experiments on three reward benchmarks demonstrate that FastGuide is up to $4.4\times$ faster than sequential reward-guided decoding while retaining similar generation quality.
发表机构
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。