从位置风险到块生存:扩散语言模型的更快生成
From Position Risks to Block Survival: Faster Generation for Diffusion Language Models
查看机构详情
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对扩散语言模型并行预测与生成不匹配的问题,提出BRISK-DLM框架,通过风险-奖励加权训练和前缀条件校正器优化提议,提升验证进展,实现高达37.4%的吞吐量提升且保持质量。
中文摘要 AI 辅助
扩散语言模型(DLM)通过并行预测多个令牌来加速生成,但这些令牌的预测方式与它们最终对生成的贡献方式之间存在不匹配。并行预测很难基于同一块内先前选定的令牌进行条件化,尽管其有效性依赖于已实现的先前部分。在流行的提议-验证解码下,这种不匹配使得错误高度不对称:早期拒绝会阻止所有后续提议对解码进展做出贡献。我们引入了BRISK-DLM,一个通过优化提议学习和选择以促进验证进展来解决这两种不匹配的框架。BRISK-DLM在自生成序列上训练,使用风险-奖励加权根据位置对验证进展和解码成本的影响动态优先排序。在推理期间,一个轻量级的基于前缀的条件校正器使用先前选择的令牌和从模型自身验证器蒸馏出的偏好对现有候选进行重新排序。该校正器重用骨干网络的并行表示,无需额外的骨干网络评估,而融合执行保持其开销较小。BRISK-DLM将端到端吞吐量提升高达37.4%,同时保持任务质量,为DLM生成建立了新的质量-吞吐量前沿。
英文摘要
Diffusion language models (DLMs) can accelerate generation by predicting multiple tokens in parallel, but there is a mismatch between how these tokens are predicted and how they ultimately contribute to generation. Parallel predictions can hardly condition on the tokens selected earlier within the same block, even though their validity depends on this realized prefix. Under the popular proposal-verification decoding, this mismatch makes errors highly asymmetric: an early rejection prevents all subsequent proposals from contributing decoding progress. We introduce BRISK-DLM, a framework that addresses both mismatches by optimizing proposal learning and selection for verified progress. BRISK-DLM trains on self-generated sequences, using risk-reward weighting to dynamically prioritize positions by their impact on verified progress and decoding cost. During inference, a lightweight prefix-conditioned corrector reranks existing candidates using previously selected tokens and preferences distilled from the model's own verifier. The corrector reuses the backbone's parallel representations and requires no additional backbone evaluation, while fused execution keeps its overhead small. BRISK-DLM improves end-to-end throughput by up to 37.4% while preserving task quality, establishing a new quality-throughput frontier for DLM generation.