arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Abra:扩散图像训练的规模化

Abra: Scaling Diffusion Image Training

Kyle Chickering, Wei-An Lin, Swayam Bhanded, Dan Saunders, Akshat Tripathi, Jiaming Song, Shyamal Buch, Xinchen Yan

arXiv 2608.17286首次发表:更新:

AI 中文总结

本研究针对文本到图像扩散模型开展系统缩放定律研究,提出Abra模型,发现其缩放规律可预测且需更多数据,相关规律可延伸至生成质量等多方面指标。

AI 中文摘要

计算最优缩放定律指导前沿语言模型的训练,但在视觉生成领域仍未得到充分探索。我们针对文本到图像的扩散模型开展了系统的缩放定律研究,使用Abra(一种受控的流匹配Transformer系列)进行训练,计算量跨度达三个数量级($10^{19}$至$10^{22}$ FLOPs),远超以往研究的计算预算。研究表明,扩散模型的缩放规律与语言模型一样可预测,但优化训练所需的数据量大得多:计算最优性出现在每个参数约200个图像标记处,是大语言模型(LLM)的Chinchilla计算最优规则的十倍。我们还发现,与语言模型不同,扩散模型对过训练具有鲁棒性,从业者应优先选择更多数据而非更大的模型。最后,我们证明这种可预测性不仅适用于训练损失,还延伸至生成质量指标、最优CFG设置、表示质量,甚至训练曲线的形态,这些曲线均收敛为通用形式。

英文摘要

Compute-optimal scaling laws guide the training of frontier language models yet remain largely unexplored for visual generation. We present a systematic scaling law study for text-to-image diffusion models using Abra, a controlled family of flow-matching transformers trained across three orders of magnitude worth of compute ($10^{19}$ to $10^{22}$ FLOPs), reaching significantly larger compute budgets than previous works. We demonstrate that diffusion models scale just as predictably as language models but require far more data to train optimally: compute optimality occurs at approximately $200$ image tokens per parameter, ten times the Chinchilla compute-optimal prescription for LLMs. We show that unlike language models, diffusion models are robust to overtraining and that practitioners should err on the side of more data rather than a larger model. Finally, we show that this predictability extends beyond training loss to generative quality metrics, optimal CFG settings, representation quality, and even the shape of the training curves, which collapse onto a universal form.

Comments25 pages, 19 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑