发表机构
Nanjing University; Nanyang Technological University; Imperial College London(南京大学; 南洋理工大学; 伦敦帝国学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AURORA-LM是一种分离文本表示构建与分布建模的连续潜在扩散语言模型,通过特定技术优化后,在OpenWebText自由生成和XSum摘要任务中性能优于同类模型。
AI 中文摘要
在生成建模领域,语言仍是异类:图像、视频和音频日益在连续潜在空间中建模,而文本生成仍主要依赖离散标记。现有连续语言模型要么继承了未针对联合生成与解码设计的嵌入空间,要么压缩自动编码的潜在变量以简化扩散过程,牺牲了标记级保真度。我们未为适配生成模型而简化表示,而是保留高容量、可解码的文本潜在变量,并设计扩散模型直接学习其分布。我们提出AURORA-LM,这是一种连续潜在扩散语言模型,将可解码文本表示的构建与其分布建模分离。基于查询的编码器-解码器将文本组织为高容量、前缀对齐的潜在序列,块因果扩散Transformer通过流匹配学习其分布,从左到右生成块,同时并行去噪每个块内的位置。由于此类潜在变量更难被扩散模型建模,AURORA-LM仅限制噪声输入通路,保留完整的干净潜在变量预测目标,在不降低解码器面向容量的同时适配全宽度潜在变量。我们进一步将噪声水平分布校准至潜在宽度,并引入自轨迹一致性以衔接独立采样的训练噪声与推理时的迭代去噪。AURORA-LM在OpenWebText自由生成和XSum摘要任务中,在评估的连续模型和基于扩散的语言模型中实现了最强性能。扩展至10亿参数、总计算量约1500 EFLOP时,性能进一步提升,在匹配的评估协议下超过了一个更大的公开潜在扩散语言模型。所有实验均在Ascend NPUs上开展。
英文摘要
Language remains an outlier in generative modeling: while images, video, and audio are increasingly modeled in continuous latent spaces, text generation still relies predominantly on discrete tokens. Existing continuous language models either inherit embedding spaces not designed for joint generation and decoding, or compress autoencoded latents to ease diffusion, sacrificing token-level fidelity. Instead of simplifying the representation to suit the generative model, we preserve a high-capacity, decodable text latent and design the diffusion model to learn its distribution directly. We introduce AURORA-LM, a continuous-latent diffusion language model that separates the construction of a decodable text representation from the modeling of its distribution. A Query-based Encoder-Decoder organizes text into a high-capacity, prefix-aligned latent sequence, and a Block-causal Diffusion Transformer learns its distribution through flow matching, generating blocks left to right while denoising positions within each block in parallel. Because such a latent is harder for diffusion to model, AURORA-LM restricts only the noisy-input pathway while retaining the full clean-latent prediction target, accommodating full-width latents without reducing decoder-facing capacity. We further calibrate the noise-level distribution to the latent width, and introduce self-trajectory consistency to bridge independently sampled training noise and iterative denoising at inference. AURORA-LM achieves the strongest performance among evaluated continuous and diffusion-based language models on OpenWebText free generation and XSum summarization. Scaling to 1B parameters with about 1500 EFLOPs of total compute yields further gains, surpassing a larger publicly released latent-diffusion language model under a matched evaluation protocol. All experiments are conducted on Ascend NPUs.
Comments40 pages, 17tables, project page: https://aurora-lm-project.github.io/