arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAI-DLLM:面向扩散语言模型的感知收敛推理

CAI-DLLM: Convergence Aware Inference for Diffusion Language Models

Farhana Amin, Sabiha Afroz, Dimitrios S. Nikolopoulos

arXiv 2608.22646首次发表:更新:

发表机构

Virginia Tech(弗吉尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CAI-DLLM 是一种无需训练的扩散语言模型推理方法,通过利用第一步置信度调整解码策略,在多类任务上实现显著推理加速,同时保持较高准确率并降低能耗。

AI 中文摘要

扩散语言模型可并行生成多个 token,但推理过程仍需重复去噪步骤,这使得生成成本较高,尤其是当模型不断重新计算已稳定的 token 时。为解决这些局限,我们提出 CAI-DLLM,一种无需训练的推理方法,它利用第一步置信度引导去噪以减少推理时间。具体而言,CAI-DLLM 会更早确定简单 token,为更难的 token 分配更多去噪步骤,并调整各输出块的解码调度。由于它仅依赖第一步置信度信号,无需重新训练、额外预测器或权重更新。我们在 LLaDA-8B-Instruct 和 Dream-7B-Instruct 上,针对数学、代码、推理、常识及长上下文任务评估 CAI-DLLM。在 LLaDA GSM8K 上,CAI-DLLM 实现最高 18.2 倍的墙钟推理加速,同时准确率从 76.27% 提升至 77.41%;在 Dream HumanEval 上实现最高 13.1 倍加速,且 pass@1 比无缓存推理更高,分别为 48.17% 与 46.95%。在更难的推理任务上,加速达 44.8 倍,最大准确率下降为 4.4 个百分点,能耗降低最高达 95.3%。

英文摘要

Diffusion language models can generate many tokens in parallel, but they still require repeated denoising steps during inference. This makes generation costly, especially when the model continues to recompute tokens that are already stable. To address these limitations, we propose CAI-DLLM, a training-free inference method that uses first-step confidence to guide denoising and reduce inference time. Specifically, CAI-DLLM commits easy tokens earlier, allocates more denoising steps to harder tokens, and adjusts decoding schedules across output blocks. As it relies only on first-step confidence signals, it does not require retraining, extra predictors, or weight updates. We evaluate CAI-DLLM on LLaDA-8B-Instruct and Dream-7B-Instruct across math, code, reasoning, commonsense, and long-context tasks. CAI-DLLM achieves up to 18.2x wall clock inference speedup on LLaDA GSM8K while improving accuracy from 76.27% to 77.41%, and up to 13.1x speedup on Dream HumanEval while achieving higher pass@1 than no-cache inference, 48.17% compared with 46.95%. On harder reasoning tasks, speedups reach 44.8x, with a largest accuracy drop of 4.4 points, while energy consumption is reduced by up to 95.3%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑