发表机构
ELLIS Institute Finland; Aalto University(芬兰ELLIS研究所; 阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对漂移模型在像素空间性能不佳的问题,提出持续表示学习以学习判别几何,并通过梯度等价性控制漂移速度,将FID降低82-95%。
AI 中文摘要
最近提出的漂移模型将迭代分布细化从推理阶段转移到训练阶段,从而实现有效的一步生成。然而,它们在复杂图像数据集上的性能强烈依赖于用于构建漂移场的表示:像素空间漂移表现不佳,而预训练特征空间则显著提高样本质量,其原因尚不清楚。我们将这一差距追溯到表示的判别几何,它决定了核密度估计(KDE)中的样本加权,进而影响漂移。我们引入了持续表示学习,它在生成器跨批次演变时持续学习更具判别性的表示几何。我们进一步建立了在匹配条件下KDE比率损失与漂移回归损失之间的当前步梯度等价性,将基于密度比的生成器优化与经验漂移联系起来,并激励对漂移速度的直接控制。在多个数据集上,我们的方法直接从像素学习有效的判别表示,并将FID相对于原始像素空间漂移模型降低约82-95%,无需预训练编码器。适配预训练表示和应用速度裁剪提供了进一步的增益。
英文摘要
Recently proposed Drifting Models shift iterative distribution refinement from inference to training, enabling effective one-step generation. However, their performance on complex image datasets depends strongly on the representation used to construct the drifting field: pixel-space drifting performs poorly, whereas pretrained feature spaces substantially improve sample quality for reasons that remain unclear. We trace this gap to the discriminative geometry of the representation, which determines sample weighting in kernel density estimation (KDE) and, consequently drift. We introduce persistent representation learning, which continuously learns a more discriminative representation geometry as the generator evolves across batches. We further establish a current-step gradient equivalence between the KDE ratio loss and drift regression loss under matched conditions, connecting density-ratio-based generator optimization to empirical drifting and motivating direct control of the drifting velocity. Across multiple datasets, our method learns effective discriminative representations directly from pixels and reduces FID by approximately $82-95\%$ over the original pixel-space Drifting Models, without pretrained encoders. Adapting pretrained representations and applying velocity clipping provide further gains.