arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩散模型的优化器:一项受控基准测试

Optimizers for Diffusion Models: A Controlled Benchmark

Arman Bolatov, Egor Shulgin, David Li, Abduragim Shtanchaev, Sebastian U. Stich, Maxim Panov, Eric Moulines, Peter Richtárik, Martin Takáč

arXiv 2609.23055首次发表:更新:

发表机构

CISPA Helmholtz Center for Information Security; KAUST; MBZUAI; EPITA(CISPA赫尔姆霍茨信息安全中心; 阿卜杜拉国王科技大学; 穆罕默德·本·扎耶德人工智能大学; 法国EPITA工程师学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次为离散扩散模型建立受控优化器基准,比较七种优化器在四种扩散公式上的表现,发现AdamW并非总是最优,且自回归预训练验证的优化器迁移良好。

AI 中文摘要

离散扩散模型现已在多个基准测试中与自回归语言模型相媲美,然而关于如何最佳训练它们的问题却受到的关注要少得多:优化器从一篇论文继承到下一篇,从未被比较过。与此同时,新的优化器几乎只在自回归预训练上进行验证,而这是在不同损失曲面上的不同目标。我们针对四种扩散公式提出了一个受控优化器基准测试,据我们所知,这是首个针对离散扩散的此类基准:在掩码扩散(text8)、均匀扩散(QM9,以及通过高斯对偶性的LM1B)和图像上的高斯扩散(CelebA-64)上,测试了七个优化器(AdamW、Lion、Muon、SOAP、MARS、MARS-M、Schedule-Free),每个任务都有已发表的参考值。每个优化器都采用相同的搜索协议,每个获胜者都在完整预算下使用三个随机种子重新训练。AdamW是一个强大的默认选择,但并不总是正确的选择:它在四个任务中的两个上被明显差距击败,且获胜者随公式而变化,因此优化器应得到与训练配方其余部分同等的关注。值得注意的是,在自回归语言模型预训练上验证的方法迁移良好:Muon、MARS-M和SOAP各自在至少一种扩散公式上击败了调优后的AdamW。该基准测试、所有运行和每个图表都可以从发布的代码端到端复现,代码见该https URL。

英文摘要

Discrete diffusion models now match autoregressive language models on several benchmarks, while the question of how best to train them has received far less attention: the optimizer is inherited from one paper to the next and never compared. New optimizers, meanwhile, are validated almost exclusively on autoregressive pretraining, a different objective on a different loss surface. We present a controlled optimizer benchmark across four diffusion formulations, to our knowledge the first for discrete diffusion: seven optimizers (AdamW, Lion, Muon, SOAP, MARS, MARS-M, Schedule-Free) on masked diffusion (text8), uniform diffusion (QM9, and LM1B through the Gaussian duality) and Gaussian diffusion on images (CelebA-64), each on a task with published reference values. Every optimizer receives the same search protocol, and every winner is retrained at the full budget with three seeds. AdamW is a strong default but not always the right choice: it is beaten by a resolved margin on two of the four tasks, and the winner changes with the formulation, so the optimizer deserves the same care as the rest of the training recipe. Notably, methods validated on autoregressive language model pretraining transfer well: Muon, MARS-M and SOAP each beat the tuned AdamW on at least one diffusion formulation. The benchmark, all runs and every figure are reproducible end to end from the released code at https://github.com/armanbolatov/diffusion-baselines.

Comments5 figures, 10 tables. Code: https://github.com/armanbolatov/diffusion-baselines

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑