简单扩散语言模型作为少步生成器的效果优于已有报道
Simple Diffusion Language Models Are More Effective Few-Step Generators Than Reported
浏览论文内容
中文总结 AI 辅助
本文证明扩散语言模型少步生成质量差距源于采样器配置不佳,通过锐化采样无需重训练即可提升性能,并引入GroupEval评估方法,强调生成器应作为评估对象。
中文摘要 AI 辅助
扩散语言模型(DLMs)承诺实现快速并行生成,然而高质量的样本往往需要大量的精化步骤,这在实际应用中削弱了其优势。这引发了对于有效少步生成新方法的广泛兴趣和快速发展。我们表明,在少步情况下,所谓的质量差距很大程度上可能源于采样器配置不佳。适度的采样器锐化,无需任何模型重训练,就能使一个几年前的掩码扩散语言模型与号称大幅改进的后继者相媲美。这种采用不同采样方式的模型,仅用16步就获得了比其标准采样器在1024步时更低的生成困惑度,同时在评判质量和语义多样性方面均有提升。我们进一步表明,传统的逐输出指标可能从根本上掩盖这些改进,因为两个此类指标之间的任何最优权衡都可以由一个仅支持最多两个输出的生成器实现。随后,我们引入了GroupEval,它分别评估质量和跨输出的语义多样性,并提供了新的见解,包括揭示蒸馏模型1.5-4.7倍的困惑度提升并未带来相应的质量提升。最后,我们解释了锐化为何有效:并行去掩码破坏了同时生成的标记之间的依赖关系,从而在预测和生成之间产生了差距。我们证明,即使在精确去噪器下,并行采样中普遍采用的温度设为1的选择通常也是次优的,且更差的预测可能产生更好的样本。通过这些结果,我们主张一个更广泛的评估原则:将部署的生成器作为比较对象,与调优的基线进行基准测试,并联合使用更符合人类对齐的度量来评估质量和多样性。
英文摘要
Diffusion language models (DLMs) promise fast parallel generation, yet high-quality samples often require large number of refinement steps, which diminishes their advantage in practice. This has led to massive interest in and rapid development of new methods for effective few-step generation. We show that much of the supposed quality gap at few steps can instead arise from a suboptimally configured sampler. Modest sampler sharpening, without any model retraining, enables a couple years old masked DLM to rival supposedly far improved successors. This differently sampled DLM in fact achieves lower generative perplexity in just 16 steps than what its standard sampler obtains with 1024, while improving both judged quality and semantic diversity. We further show that conventional per-output metrics can fundamentally obscure these gains, since any optimal trade-off between two such metrics can be attained by a generator supported on at most two outputs. We subsequently introduce GroupEval, which separately evaluates quality and across-output semantic diversity, and offers fresh insights including uncovering how 1.5-4.7x perplexity gains of a distilled model yield no corresponding quality gain. Finally, we explain why sharpening helps: parallel unmasking destroys dependencies among simultaneously generated tokens, creating a gap between prediction and generation. We prove that pervasive temperature choice of one is generically suboptimal under parallel sampling even for an exact denoiser, and that worse predictions can yield better samples. Through these results, we argue for a broader evaluation principle of treating the deployed generator as the object of comparison, benchmarking it against tuned baselines, and assessing quality and diversity jointly and with more human-aligned measures.
发表机构
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。