大型语言连续扩散模型
Large Language Continuous Diffusion Models
浏览论文内容
中文总结 AI 辅助
提出首个大规模连续扩散语言模型Sigma,通过低维ODE/SDE轨迹和分块似然训练,实现与离散模型相当的性能,并具备引导质量-多样性权衡和高效蒸馏等独特优势。
中文摘要 AI 辅助
尽管离散扩散语言模型(dLMs)在快速并行解码方面取得了成功,但其非平滑、高维空间阻碍了用于推理和推理加速的轨迹引导。为了克服这一问题,我们提出了Sigma,这是首个基于可引导、低维ODE/SDE潜在轨迹的大规模(3B/8B)连续dLM。通过似然优化的分块训练,Sigma在联合去噪高斯损坏的令牌嵌入的同时学习最优嵌入几何。为了加速训练,Sigma利用自回归(AR)模型的预训练权重进行热启动。在推理过程中,我们确定无分类器引导和分数温度对于高保真推理和编码至关重要。在与最先进的离散对应模型(掩码dLMs和AR基线)进行的全面数学推理和编码评估中,Sigma在预训练后的标准基准(如GSM8K、Minerva、HumanEval、MBPP)上取得了与离散模型相当的性能,并在监督微调后的挑战性推理任务(如MATH-500、AIME)上同样如此。除了性能持平外,我们揭示了连续dLMs独有的关键结构特性:(i)嵌入空间引导有效控制质量-多样性权衡,产生强大的pass@k性能;(ii)连续轨迹使得低NFE下优雅降级和高效蒸馏成为可能。这些确立了连续dLMs作为高效语言生成的有前景范式的地位。
英文摘要
Despite the success of discrete diffusion language models (dLMs) for fast parallel decoding, their non-smooth, high-dimensional space hinders trajectory steering for reasoning and inference acceleration. To overcome this, we present Sigma, the first large-scale (3B/8B) continuous dLM built on steerable, low-dimensional ODE/SDE latent trajectories. Trained blockwise via likelihood optimization, Sigma jointly denoises Gaussian-corrupted token embeddings while learning an optimal embedding geometry. To accelerate training, Sigma leverages pre-trained weights from autoregressive (AR) models for warm-starting. During inference, we identify classifier-free guidance and score temperature as essential for high-fidelity reasoning and coding. Across comprehensive math reasoning and coding evaluations against state-of-the-art discrete counterparts (masked dLMs and AR baselines), Sigma achieves competitive performance with discrete models on standard benchmarks (e.g., GSM8K, Minerva, HumanEval, MBPP) after pre-training and on challenging reasoning tasks (e.g., MATH-500, AIME) after supervised fine-tuning. Beyond performance parity, we uncover key structural properties unique to continuous dLMs: (i) embedding-space steering effectively governs the quality-diversity trade-off, yielding strong pass@k performance and (ii) continuous trajectories enable graceful degradation for low NFEs and efficient distillation. These establish continuous dLMs as a promising paradigm for efficient language generation.
发表机构
- NVIDIA(英伟达)
- Cornell University(康奈尔大学)
- Georgia Institute of Technology(佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。