arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

时间锚定扩散语言模型:用于快速生成的潜空间缓存

Time-Anchored Diffusion Language Models: Latent-Space Caching for Fast Generation

Joel Anto Paul, Litu Rout, Aditya Akella, Sanjay Shakkottai

arXiv 2609.37924首次发表:更新:

发表机构

UT Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出时间锚定扩散语言模型,通过潜空间缓存复用锚点加速生成,无需监督目标,在DiffusionGemma-26B上吞吐量提升最高79%,计算减少38%。

AI 中文摘要

近期关于锚定扩散语言模型的研究通过使用有监督的重要词元目标塑造中间潜空间来改进去噪过程。在本工作中,我们引入了基于时间的(自监督)锚定方法,该方法无需此类目标即可学习并复用潜锚点。我们的关键观察是,锚点编码了干净序列的持久属性,例如其语义意图、全局结构或中间计划。尽管随着词元画布的演变,它们的隐藏表示会变得过时,但其语义内容在邻近的扩散时间上仍然有用。这通过一个两阶段架构实现,该架构由一个相对昂贵的锚点网络(生成潜缓存状态)和一个轻量级去噪网络(在每个反向步骤中通过融合模块将缓存的潜状态与当前状态智能地组合)组成。这赋予了锚定一种潜空间缓存解释:锚点网络被周期性评估,而其缓存表示在多个反向步骤中被复用。我们将此框架实例化为TADM:Post-train,它对预训练的DLM进行时间锚定,以及TADM:Pretraining,它在预训练期间学习基于时间的锚点。应用于DiffusionGemma-26B时,TADM:Post-train在多个数学、代码和STEM基准(GSM8K、AIME26、GPQA-Diamond、LiveCodeBench-v6、HumanEval、MMLU-Pro)上将吞吐量提高了约49%至79%。TADM:Pretraining相对于标准的单阶段DLM将Transformer层计算减少了高达38%,并且比ADLM实现了高达73%的更高实测吞吐量。

英文摘要

Recent work on anchored diffusion language models improves denoising by shaping an intermediate latent space with supervised important-token targets. In this work, we introduce time-based (self-supervised) anchoring, which learns and reuses latent anchors without requiring such targets. Our key observation is that anchors encode persistent properties of the clean sequence, such as its semantic intent, global structure, or intermediate plan. Although their hidden representations become stale as the token canvas evolves, their semantic content remains useful across nearby diffusion times. This is implemented through a two-stage architecture consisting of a relatively expensive anchor network that generates the latent cache state and a lightweight denoising network that intelligently combines the cached latent state with the current state at each reverse step using a fusion module. This gives anchoring a latent-space caching interpretation: the anchor network is evaluated periodically, while its cached representation is reused across multiple reverse steps. We instantiate this framework as TADM:Post-train, which time-anchorizes pretrained DLMs, and TADM:Pretraining, which learns time-based anchors during pretraining. Applied to DiffusionGemma-26B, TADM:Post-train improves throughput by approximately 49% to 79% on several math, code, and STEM benchmarks (GSM8K, AIME26, GPQA-Diamond, LiveCodeBench-v6, HumanEval, MMLU-Pro). TADM:Pretraining reduces Transformer-layer computation by up to 38% relative to a standard single-stage DLM, achieves up to 73% higher measured throughput than ADLM.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑