arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从原子到熵:凸域中扩散训练的最优噪声分配

From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime

Luca Ambrogioni, Giulio Franzese, Alberto Foresti, Gabriel Raya, Bac Nguyen, Georgios Batzolis, Yuhta Takida, Naoki Murata, Chieh-Hsin Lai, Yuki Mitsufuji

arXiv 2607.20540首次发表:更新:

发表机构

Radboud University; EURECOM; Tilburg University; Sony AI; University of Cambridge; Sony Group Corporation(拉德堡德大学; 欧洲电信系统研究所; 蒂尔堡大学; 索尼人工智能公司; 剑桥大学; 索尼集团公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究扩散模型训练时噪声水平分配问题,开发通用统计框架,在不同 regime 下得出最优分配结论,并通过实验验证,平方根熵调度能提升离散域训练效率且在连续图像上有竞争力。

AI 中文摘要

扩散模型应如何决定在哪些噪声水平上进行训练以及训练量多少?当前噪声调度大多基于启发式或经验调整。本文开发了一个通用统计框架来研究扩散训练中渐近最优噪声水平分配。在完全耦合 regime 下,在凸性等假设下,优化训练调度有原子极小值,集中在有限多个噪声水平。在理想化独立学习者 regime 下,通过随机矩阵分析得出信息论代理。在可控设置中测试了这些预测,结果表明平方根熵调度可提高离散域训练效率,在连续图像上与标准启发式竞争。

英文摘要

How should a diffusion model decide which noise levels to train on, and how much? Despite the importance of this choice, current noise schedules are based largely on heuristics or empirical tuning. Here, we develop a general statistical framework for studying asymptotically optimal noise-level allocation in diffusion training. Our first main result concerns the fully coupled regime, where information can spread between different time points. Under convexity or Polyak-Lojasiewicz-type assumptions, we show that the optimized training schedule admits an atomic minimizer, concentrated on finitely many noise levels. Our second main result specializes this framework to an idealized independent-learner regime, intended to model temporal specialization in neural networks. Under an additional feature-noise decoupling condition, a random-matrix analysis leads to an information-theoretic proxy: the decoupled sampling density is proportional to the square root of the generative entropy rate, the rate at which conditional entropy grows along the forward process. We test these predictions in controlled settings where the coupled objective can be optimized directly, including Dirac mixtures, low-dimensional manifolds, and MNIST. In these settings, the optimized schedules are consistently finite-support, while the smooth entropic proxy closely tracks the atomic optimum in neural-network models and breaks down mainly in the fully coupled parametric case, as the theory suggests. We then evaluate the entropic schedule in larger-scale experiments, where full schedule optimization is currently intractable. The results indicate that square-root entropy scheduling can substantially improve training efficiency on discrete domains and remains competitive with standard EDM-style heuristics on continuous images.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑