arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分层连续扩散语言模型

Hierarchical Continuous Diffusion Language Models

Hui Ren, Zihan Li, Chang Liu, Huidong Liu, Alexander Schwing

arXiv 2610.02193首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Amazon.com, Inc.(伊利诺伊大学厄巴纳-尚佩恩分校; 亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出分层连续扩散语言模型(HC-DLM),通过将离散标记生成与连续潜在轨迹耦合于统一去噪过程,解决扩散模型并行解码中标记独立采样及状态与标记配置脱节的问题,在数独、Countdown和LM1B上超越离散与连续扩散基线。

AI 中文摘要

离散扩散语言模型为需要双向推理和全局约束满足的任务提供了一种引人注目的自回归生成替代方案。然而,它们共享一个结构性瓶颈:在并行解码时,每个标记独立地从其边缘分布中采样,切断了同时解码的标记之间的统计依赖性。连续扩散语言模型通过去噪共享的连续状态来避免这一问题,但其去噪器仅能看到该状态,因此在最终解码之前,没有任何东西将其与有效的标记配置联系起来。为了解决这一问题,我们提出了分层连续扩散语言模型(HC-DLM),该模型在单一原则性的去噪过程中将离散标记生成与连续潜在轨迹耦合,其训练目标源自标记似然的变分下界。与最近将连续上下文附加到自包含离散链的方法不同,HC-DLM使潜在状态成为唯一持久的生成状态:每一步都从中读出标记,并将其反馈为下一次潜在更新的脚手架。在结构化推理(数独)、数学规划(Countdown)和语言建模(LM1B)上,HC-DLM在匹配模型大小的情况下,在数独和Countdown的谜题准确率以及LM1B的生成困惑度上均优于离散和连续扩散基线。项目页面:此 https URL。

英文摘要

Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update. On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B. Project page: https://hc-dlm.github.io/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑