arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15655cs.CLcs.LG

扩散语言模型的自适应多步前瞻解码

Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models

Yingqian Cui, Wei Deng, Lantao Mei, Hang Li, Charu C. Aggarwal, Hui Liu, Yue Xing

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对扩散语言模型解码,提出自适应多步前瞻框架AdaLook,基于候选分数方差动态决定是否继续展开及进行分支扩展,避免不必要的深度展开,实验证明其在准确性和解码步骤权衡上优于现有一步前瞻解码方法。

中文摘要 AI 辅助

掩码扩散语言模型(DLMs)通过迭代细化掩码令牌实现并行文本生成,为自回归解码提供了有前景的替代方案。近期基于前瞻的解码方法通过探索未来解码状态改善了准确性与效率的权衡。但现有方法主要依赖浅层次的一步前瞻,对更长解码轨迹并非最优。我们发现深度前瞻的简单扩展也无效。因此,本文提出AdaLook,一个用于DLM解码的自适应前瞻框架。它基于候选分数方差动态决定是否继续展开,并在中间展开状态需要额外探索时进行分支扩展。实验表明AdaLook比现有一步前瞻解码方法在准确性和解码步骤权衡上表现更好。

英文摘要

Masked diffusion language models (DLMs) enable parallel text generation by iteratively refining masked tokens, offering a promising alternative to autoregressive decoding. Recent lookahead-based decoding methods improve the accuracy--efficiency trade-off by exploring future decoding states before committing token updates. However, existing approaches mainly rely on shallow one-step lookahead, which optimizes immediate information gain but can be suboptimal for longer-horizon decoding trajectories. Meanwhile, we find that a naive extension for deeper lookahead is also ineffective, as fixed-depth rollout introduces additional computation and cannot adapt to heterogeneous intermediate decoding states. Thus, in this work, we propose AdaLook, an adaptive lookahead framework for DLM decoding. AdaLook dynamically determines whether to continue rollout based on candidate-score variance and further enables branch expansion when intermediate rollout states require additional exploration. This design avoids unnecessary deep rollout while allowing the decoder to re-trigger lookahead from informative intermediate states. Experiments on various benchmarks and models demonstrate that AdaLook achieves a better accuracy--decoding steps trade-off than existing one-step lookahead decoding methods.

发表机构

  • Michigan State University(密歇根州立大学)
  • Morgan Stanley(摩根士丹利)
  • IBM T.J. Watson Research Center(IBM 托马斯·J·沃森研究中心)

机构由 AI 辅助整理,请以论文原文为准。

↑