DiffusionGemma 技术报告
DiffusionGemma Technical Report
浏览论文内容
中文总结 AI 辅助
本研究推出开放权重语言模型 DiffusionGemma,通过并行优化 256 个 token 块规避 AR 模型顺序解码瓶颈,经两阶段微调后实现 1500 token/秒生成速度,性能优于带推测解码的 AR 模型,还支持多模态等功能并为混合解码提供方向。
中文摘要 AI 辅助
我们推出 DiffusionGemma,这是一款采用离散扩散技术生成文本的实验性开放权重语言模型,具备极高的生成速度。与传统自回归(AR)大型语言模型逐 token 解码的方式不同,DiffusionGemma 会并行迭代优化 256 个 token 的块,从而规避了传统 AR 模型的顺序解码瓶颈。我们并非从零开始训练 DiffusionGemma,而是通过对混合专家模型 Gemma 4 进行微调得到该模型,Gemma 4 拥有 38 亿激活参数和 252 亿总参数。我们的计算高效的两阶段训练流程使用了不到原 AR 模型总训练 token 预算的 10%:第一阶段采用监督微调来学习双向去噪;第二阶段将强化学习与采样器蒸馏相结合,共同提升生成质量和推理效率。DiffusionGemma 在生成速度与模型能力的权衡上建立了新的帕累托前沿。在完整评估套件的平均表现中,它每次前向传递生成约 20 个 token,在单个 NVIDIA H100 GPU 上实现了约 1500 个输出 token 每秒的速度,即便与采用最先进推测解码的 AR 模型相比也快得多。DiffusionGemma 还保留了原模型对思考模式、多模态输入和长上下文的支持。尽管经过了扩散微调,它仍具备 AR 生成能力,且性能仅出现轻微下降,这为混合扩散-AR 解码指明了方向。
英文摘要
We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at a time, DiffusionGemma iteratively refines blocks of 256 tokens in parallel, avoiding the sequential decoding bottleneck of conventional autoregressive (AR) large language models. Instead of training from scratch, we obtain DiffusionGemma by fine-tuning the mixture-of-experts Gemma 4 model with 3.8B activated and 25.2B total parameters. Our compute-efficient two-stage training pipeline uses fewer than 10% of the starting AR model's total training token budget. The first stage uses supervised fine-tuning to teach bidirectional denoising, while the second stage combines reinforcement learning with sampler distillation to jointly improve generation quality and inference efficiency. DiffusionGemma establishes a new Pareto frontier for the trade-off between generation speed and model capability. Averaged across our full evaluation suite, it generates around 20 tokens per forward pass and achieves roughly 1,500 output tokens per second on a single NVIDIA H100 GPU, which is substantially faster than AR models even with state-of-the-art speculative decoding. DiffusionGemma also retains the starting model's support for thinking mode, multimodal inputs, and long contexts. Despite diffusion fine-tuning, it remains capable of AR generation with only minor performance degradation, suggesting a path toward hybrid diffusion-AR decoding.