AI 中文总结
研究提出令牌时间连续扩散(TTCD)语言模型,其在连续空间运行,引入令牌时间概念。该模型避免多令牌并行采样,能更好建模条件生成。实验表明,在高速加速时TTCD优于离散模型,在无条件和条件生成上有优势,数独求解也有类似成果。
AI 中文摘要
本文介绍了令牌时间连续扩散(TTCD),一种新的扩散语言模型。它在连续空间中运行,将高斯噪声确定性地映射到最终令牌画布,无需进一步采样。关键的是,它引入了每个令牌时间的新概念,一些令牌从噪声到令牌的速度比其他令牌快。连续空间建模有助于TTCD避免多个令牌的并行采样,这是纯离散空间迭代模型在高速加速时不准确的关键来源。每个令牌时间的概念有助于TTCD更好地对条件生成进行建模,允许更确定的令牌以更快的速度进行,并在细化过程中允许差异化的令牌间影响。在高速加速时,TTCD优于离散模型。我们在OpenWebText上训练了一个1.6亿参数的TTCD模型,然后进行自蒸馏;发现在高速加速时,我们在无条件生成质量上具有可比性,在条件生成方面优于几个在相同数据上训练和自蒸馏的类似规模的现有模型。在数独求解中也取得了类似的收益。
英文摘要
In this paper we introduce token time continuous diffusion (TTCD), a new diffusion language model which (a) operates in continuous space, deterministically mapping Gaussian noise to a final token canvas with no further sampling, and crucially (b) incorporates a new notion of per-token times, with some tokens proceeding from noise to token at a faster rate than others. Continuous space modeling helps TTCD avoid the parallel sampling of multiple tokens, which is a key source of inaccuracy at high speedups for models that iterate purely in discrete space. The notion of per-token times helps TTCD to better model conditional generation, allows for more sure tokens to proceed at a faster rate, and allows for differentiated inter-token influences during refinement. TTCD outperforms discrete models at high speedups. We train a 160M parameter TTCD model on OpenWebText, and then self-distill it; we find that at high speedups we are comparable in unconditional generation quality, and outperform in conditional generation, several existing models of similar size trained, on the same data, and self-distilled. We achieve similar gains in Sudoku solving as well.