在降采样图像上进行训练何时能产生相同的梯度?
When does training on downscaled images yield the same gradients?
AI总结:
该研究探究降采样图像训练的梯度一致性,推导梯度变化的两个项,发现特定噪声窗口内降采样梯度与原生梯度接近,据此训练LoRA适配器可减少14.6%训练时间且性能接近原生。
AI中文摘要:
扩散Transformer在图像生成领域表现出色,但其训练成本随分辨率呈超线性增长。近期研究基于频谱前提论证了在降低分辨率下进行训练或采样的合理性:在高噪声环境中,降采样的隐变量几乎保留了全部剩余信号。然而,降采样步骤是否也能保留原生训练梯度信号,这一问题仍未得到解决。我们将该信号在降采样下的变化简化为两个项:一项是由降采样率决定的噪声相关项,在高噪声下会如频谱前提所预测的那样衰减;另一项是由目标网格的绝对token数决定的与σ无关的基底,该基底由计算图本身携带,任何噪声水平都无法将其消除。所测得的(路径,σ)图证实了这一解释,并揭示了频谱图景无法表达的结构:在1024→768路径上,存在一个窗口(0.65 < σ < 0.95),该窗口在任何容差下都没有频谱准则可预测,在此窗口内,降采样梯度与原生梯度的差值保持在很小的范围内。在固定步数预算下,将LoRA适配器的降采样步骤限制在该图验证的路径和噪声窗口内进行训练,可将训练时间减少14.6%,同时在权重空间中仍接近原生状态。代码可在该https URL获取。
英文摘要:
Diffusion transformers deliver strong image generation, but their training cost grows superlinearly with resolution. Recent work justifies training or sampling at reduced resolution on a spectral premise: at high noise, a downscaled latent preserves almost the full surviving signal. Whether a downscaled step also preserves the native training gradient signal, however, has remained unresolved. We reduce how that signal changes under downscaling to two terms: a noise-dependent term governed by the downscale ratio, which decays at high noise as the spectral premise predicts, and a σ-independent floor governed by the target grid's absolute token count, carried by the compute graph itself and removed by no noise level. The measured (route, σ) map corroborates the account and uncovers structure the spectral picture cannot express: on the 1024->768 route, a window (0.65 < σ< 0.95), predicted by no spectral criterion at any tolerance, where the downscaled gradient stays within a small margin of the native one. Training LoRA adapters with downscaled steps restricted to the routes and noise windows the map validates reduces training time by 14.6% at a fixed step budget while remaining near-native in weight space. Code is available at https://github.com/sorryhyun/anima_lora.