发表机构
National Yang Ming Chiao Tung University; University at Albany, SUNY(国立阳明交通大学; 纽约州立大学奥尔巴尼分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对扩散语言模型解码过早终止问题,提出无训练的候选感知提前退出框架LATCH,结合CVC与BWEC,在零样本11项任务上,以准确率损失极小的代价实现多倍解码速度提升。
AI 中文摘要
扩散语言模型(DLM)在每一步去噪时都会输出临时预测,这为生成时提前退出提供了可能,即未完成整个解码流程就停止解码。现有提前退出门控通过固定区域的置信度统计或依赖解码流程的规则来决定终止时机,但这些证据过于粗略,无法应对需一次性冻结所有剩余位置的决策,因此在长思维链输出上会过早触发,这类输出的答案仅在流程末尾才稳定。另一种无训练加速方式是自适应采样,它控制解码时位置提交的速度,但从未验证输出本身是否稳定。我们提出一种无训练、候选感知的提前退出框架,将两种加速维度分离,使每个决策匹配对应范围的证据:置信度验证提交(CVC)负责决定序列是否可停止,它通过各任务输出格式指定的确定性解析器,验证动态提取的候选跨度上的置信度与持续的argmax稳定性;块级提前提交(BWEC)负责决定加速位置,它对非最终块应用成本更低的局部规则,最终块和全局终止则由CVC控制,我们将二者的组合称为LATCH(带跟踪候选停止的局部加速)。与现有方法不同,LATCH无需构建后缀提示,它无提示锚点但具备格式感知能力。我们在零样本设置下,使用LLaDA和Dream在11项任务上对LATCH进行端到端评估。在全部22种评估设置中,LATCH的准确率与全解码的差距不超过2.0个百分点,且仅需一组冻结的超参数即可跨模型主干迁移、无需调整,同时在短答案任务上实现9.3-17.8倍的端到端每秒令牌数(TPS)加速,在长推理任务上实现2.0-3.3倍的TPS加速。
英文摘要
Diffusion language models expose a provisional prediction at every denoising step, and on many tasks the candidate answer inside it stabilizes before the step schedule is exhausted. This creates two acceleration opportunities, leaving a block early and stopping the sequence early, but the two require different criteria because block acceleration is local whereas sequence termination is global and freezes the graded answer. Existing methods usually optimize only one axis, and existing exit gates rely on fixed-region confidence or schedule-dependent rules rather than the candidate answer itself. We present $\textbf{C}^4$, which coordinates the two axes by giving each decision its own gate. $\textbf{C}$onfidence-Verified Early Exit (CVEE) decides when the sequence may stop, requiring confidence and sustained argmax stability over a candidate span re-extracted at every step. $\textbf{C}$ommit-$\textbf{C}$ore-Then-$\textbf{C}$onfirm (CCTC) decides which token positions a step may commit by borrowing an autoregressive freezing order inside each block: it commits a boundary-anchored core and confirms deferred positions one step later, so the answer block can be accelerated without allowing local commits inside the answer span to determine sequence-level termination. On 12 zero-shot tasks with LLaDA and Dream, one frozen configuration removes 64--95% of decoding steps and delivers measured end-to-end speedups of 2.6 to 8.6 over full decoding. Code is available at https://github.com/ming053l/C4-dLLM.
CommentsCode is available at https://github.com/ming053l/C4-dLLM