arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

预测而非迭代:用于扩散语言模型的高效自适应长度 infilling

Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models

Haobo Xu, Sirui Chen, Yuanchen Bei, Lingjie Chen, Yuchen Yan, Dongqi Fu, Jingrui He, Hanghang Tong

arXiv 2609.02108首次发表:更新:

发表机构

University of Illinois at Urbana-Champaign; Meta(伊利诺伊大学厄巴纳-香槟分校; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对扩散语言模型 infilling 任务的局限,提出无预设长度的高效 PILL 方法,在多模型多基准上实现代码通过率、文本 BLEU-2 提升且推理速度加快。

AI 中文摘要

扩散语言模型(DLMs)已成为自回归范式的有前景替代方案,凭借双向注意力和任意顺序生成能力,DLMs 天然适配 infilling 任务——该任务需在前缀和后缀的条件下生成中间片段。然而 infilling 对片段长度敏感,而 DLMs 要求生成前固定长度。尽管已有研究将 DLMs 扩展至动态长度,但仍存在两个局限:(i)对初始长度敏感,这些方法需预设长度以初始化搜索,且对该初始长度高度敏感,常产生次优结果;(ii)推理效率低,它们要么在生成过程中插入长度变更操作,要么通过多步去噪置信度反复搜索合适长度,两者均引入大量额外前向传播和计算成本。因此,我们提出 PILL(基于探测的无预设长度解码 infilling),这是一种用于 DLMs 的高效 infilling 方法,无需预设初始长度,且比基线添加的额外前向传播少得多,大幅缩短推理时间。实验在涵盖不同家族、架构和训练方案的 5 个 DLMs 及 8 个 infilling 基准上开展,结果显示,PILL 相比最强基线,代码的平均通过率提升 4.8,文本的 BLEU-2 提升 6.0,同时运行速度比该基线快 1.82 倍。代码可在该 https URL 获取。

英文摘要

Diffusion language models (DLMs) have emerged as a promising alternative to the auto-regressive paradigm. With bidirectional attention and any-order generation, DLMs naturally fit infilling tasks, which require generating a middle span conditioned on both the prefix and the suffix. However, infilling is sensitive to the length of the span, while DLMs require the length to be fixed before generation. Although prior studies extend DLMs to dynamic lengths, they still suffer from two limitations. (i) Sensitivity to initial length. These methods require a preset length to initialize the search and are highly sensitive to this initial length, often yielding suboptimal results. (ii) Inference inefficiency. They either insert length-changing operations during generation or repeatedly search for an appropriate length using multi-step denoising confidence, both of which introduce substantial extra forward passes and computational cost. Therefore, we propose PILL (Probing-based InfiLling with preset-Length-free decoding), an efficient infilling method for DLMs that requires no preset initial length and adds far fewer extra forward passes than baselines, substantially reducing inference time. Experiments show that, across five DLMs spanning different families, architectures, and training recipes on eight infilling benchmarks, PILL improves over the strongest baseline by +4.8 average pass rate on code and +6.0 BLEU-2 on text, while running 1.82x faster than that baseline. The code is available at https://github.com/Hsu1023/PILL.

CommentsAccepted at EMNLP 2026 (Main Conference)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑