arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CORA-Diff:面向高效扩散语言模型推理的置信导向残差接受

CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference

Yifan Wu, Yufeng Zhang, Kenli Li

arXiv 2608.11235首次发表:更新:

AI 中文总结

本文提出无需训练的CORA-Diff方法,利用原生置信度与持久性门控识别残差位置,在多基准测试中显著降低扩散语言模型推理的重复计算,同时保持任务性能。

AI 中文摘要

扩散语言模型(DLM)会并行更新多个token,但实际解码器通常采用固定去噪 horizon。许多预测会提前稳定,但分块解码会持续直到所有位置都被解析,导致重复的密集前向传播。现有加速器常依赖学习滤波器、修改后的分数、依赖模型或特定缓存机制。本文探究原生轨迹信号是否能识别可能与确定性密集端点匹配的残差位置。我们提出CORA-Diff,这是一种无需训练的方法,它保留原始转移规则,仅对规则未解决的位置应用置信度与持久性门控。被接受的token会作为上下文保持可见,且当所有位置被解析后,分块终止。这无需修改主干网络、学习接受模型或修改logit。我们的理论解释了为何高置信度、持久的预测更可能与固定horizon的密集端点匹配,配对的干预后轨迹提供了直接的实证支持。我们在单独的GSM8K校准子集上选择一个操作点,并将其冻结用于所有评估。在匹配的Learn2PD风格LLaDA协议下,CORA-Diff在全部8种任务长度设置中测得的运行时间最低。在5种设置中,任务得分与密集解码相当或更高,最大观测降幅为1.22分。它在GSM8K和HumanEval上相对于感知EOS的密集解码的增量加速比分别为2.70倍和3.32倍;在固定horizon 1024/1024机制隔离协议下达到13.14倍加速,且无需重新调优即可迁移到Dream,加速比为3.18倍至3.53倍。这些结果表明,原生置信度与持久性可实现可靠的残差接受,在保留任务质量的同时减少重复去噪计算。

英文摘要

Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon. Many predictions stabilize early, but blockwise decoding continues until all positions are resolved, causing repeated dense forward passes. Existing accelerators often rely on learned filters, modified scores, dependency models, or cache-specific mechanisms. We ask whether native trajectory signals can identify residual positions likely to match the deterministic dense endpoint. We propose CORA-Diff, a training-free method that preserves the original transfer rule and applies confidence-and-persistence gating only to positions that rule leaves unresolved. Accepted tokens remain visible as context, and the block terminates once all positions are resolved. This requires no backbone change, learned acceptance model, or logit modification. Our theory explains why high-confidence, persistent predictions are more likely to match the fixed-horizon dense endpoint, and paired post-intervention trajectories provide direct empirical support. We select one operating point on a separate GSM8K calibration subset and freeze it for all evaluations. Under a matched Learn2PD-style LLaDA protocol, CORA-Diff has the lowest measured runtime in all eight task-length settings. Task scores match or exceed dense decoding in five settings, and the largest observed drop is 1.22 points. Its incremental speedups over EOS-aware dense decoding are 2.70x and 3.32x on GSM8K and HumanEval. It also reaches 13.14x under the fixed-horizon 1024/1024 mechanism-isolation protocol and transfers to Dream without retuning at 3.18x-3.53x. These results show that native confidence and persistence enable reliable residual acceptance, reducing repeated denoising computation while preserving task quality.

Comments9 pages, 2 figures, 3 tables. Code: https://github.com/wyffffff/cora-diff-llada

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑