发表机构
School of Computer Science, Shanghai Jiao Tong University; College of Artificial Intelligence, Nankai University; School of Computer Science, Peking University(上海交通大学计算机科学学院; 南开大学人工智能学院; 北京大学计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Archer方法,通过自适应复用缓存的提示隐藏状态,在保证扩散语言模型回滚能力的同时,实现了2.57倍平均加速与33.63%最佳平均性能,还提升了Pass@1指标。
AI 中文摘要
扩散语言模型(DLMs)会迭代地优化序列,允许随着上下文演变修正早期预测,这种回滚能力是其与不可逆自回归生成的区别,但也导致推理成本高昂。每次去噪更新都会改变全局上下文,迫使提示和响应状态都需重新计算,尽管只有响应标记可被修正。键值(KV)缓存可降低此成本,但传统缓存假设历史状态不可变,因此难以与回滚能力兼容。本文提出一种无需训练的KV缓存方法——Archer(用于扩散语言模型高效回滚的缓存隐藏状态自适应复用),适用于具备回滚能力的DLMs。Archer以非对称方式使可变响应与当前假设同步,同时在有限状态邻域内复用提示的键/值。尽管提示表示在双向注意力下也会变化,但其标记身份保持固定,因此有限复用可分摊重复的提示计算,无需缓存可变响应状态。该方法还延迟了试探性标记的反馈,减少了瞬态高置信度错误的过早强化,为回滚提供更多修正机会。我们的分析将提示复用表征为与可逆性对齐的缓存边界,界定了其依赖状态的近似误差,并给出保留全刷新的解码器边际条件。扩散语言模型加速常以质量换速度,Archer则突破了这一权衡边界,在主测试集上取得了33.63%的最佳平均性能,同时实现了2.57倍的平均加速;在所有评估设置中,其Pass@1指标提升最高达3.05个百分点,加速最高达2.95倍。受控分析将质量提升与延迟提示反馈关联,并验证了感知状态的刷新机制。代码可在指定链接获取。
英文摘要
Diffusion language models (DLMs) iteratively refine a sequence, allowing earlier predictions to be revised as context evolves. This rollback capability distinguishes them from irreversible autoregressive generation, but makes inference costly. Every denoising update alters the global context, forcing both prompt and response states to be recomputed even though only response tokens are revisable. Key-value (KV) caching could reduce this cost, yet conventional caching assumes immutable historical states and is therefore difficult to reconcile with rollback. In this paper, we introduce Adaptive Reuse of Cached Hidden States for Efficient Rollback (Archer), a training-free KV caching method for rollback-capable DLMs. Archer asymmetrically keeps the mutable response synchronized with the current hypothesis while reusing prompt K/V within a bounded state neighborhood. Although prompt representations also change under bidirectional attention, their token identities remain fixed; bounded reuse therefore amortizes repeated prompt computation without caching mutable response states. It also delays feedback from tentative tokens, reducing premature reinforcement of transient high-confidence errors and giving rollback more opportunity to correct them. Our analysis characterizes prompt reuse as a reversibility-aligned cache boundary, bounds its state-dependent approximation error, and gives a decoder-margin condition for preserving full-refresh decisions. Existing DLM acceleration often trades quality for speed. Archer shifts this frontier, attaining the best mean performance of 33.63% together with a 2.57x mean speedup on the main suite. Across evaluated settings, it improves Pass@1 by up to 3.05 points and reaches up to 2.95x speedup. Controlled analyses connect the quality gain to delayed prompt feedback and validate state-aware refresh. Our code is available at https://github.com/Hxnng/Archer.