发表机构
École Polytechnique; CNRS; IP Paris; AMIAD; École Nationale des Ponts et Chaussées; UGE; UC Berkeley(巴黎综合理工学院; 法国国家科学研究中心; 巴黎理工学院; AMIAD; 巴黎路桥学院; 巴黎第八大学; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对像素空间流模型训练的频谱失衡问题,提出Focal Log-Frequency Loss(f-loss),结合频率与像素监督的训练策略,可将流匹配收敛速度提升最高40%,同时改善FID和感知保真度。
AI 中文摘要
自然图像遵循$1/f^2$频谱分布:大部分信号能量位于低空间频率,而纹理、边缘等感知重要结构则占据稀疏的高频带。然而,像素空间重建目标对所有空间误差一视同仁,导致低频率主导优化信号,延缓了精细尺度细节的学习。在本研究中,我们将这种目标层面的频谱失衡识别为像素空间流模型训练中的关键低效因素。为解决该问题,我们提出了Focal Log-Frequency Loss(f-loss),这是一种频谱平衡的目标函数,可均衡各频率的学习信号,强化像素空间目标中原本代表性不足的高频分量。在此基础上,我们引入了一种简单的训练策略,结合频率与像素监督:我们在训练初期强调频率域学习以捕获所有频率,随后过渡到标准像素空间v-loss进行空间细化。这种平衡策略缓解了像素损失的低频偏差,使训练信号与模型的演进需求相契合。我们的方法概念简单,无需修改架构,可作为流匹配损失的即插即用替代方案。在多种模型规模下,它将收敛速度提升了最高40%,同时持续改善FID和感知保真度。我们将发布代码和模型。
英文摘要
Natural images follow a $1/f^2$ spectral distribution: most signal energy lies in the low spatial frequencies, while the perceptually important structures such as textures and edges occupy sparse high-frequency bands. Pixel-space reconstruction objectives, however, treat all spatial errors uniformly, causing low frequencies to dominate the optimization signal and delaying the learning of fine-scale details. In this work, we identify this objective-level spectral imbalance as a key inefficiency in training pixel-space flow models. To address it, we propose a Focal Log-Frequency Loss (f-loss), a spectrally balanced objective that equalizes the learning signal across frequencies, emphasizing high-frequency components that are otherwise underrepresented in pixel-space objectives. Building on this, we introduce a simple training strategy that combines frequency and pixel supervision: we first emphasize frequency-domain learning early to capture all frequencies, and then transition to standard pixel-space v-loss for spatial refinement. This balancing mitigates the low-frequency bias of pixel losses and aligns the training signal with the evolving needs of the model. Our approach is conceptually simple, requires no architectural changes, and acts as a drop-in replacement for flow matching losses. Across multiple model scales, it accelerates convergence by up to 40% while consistently improving FID and perceptual fidelity. We will release code and models.