arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RepFlow:互惠监督提升流模型中的生成与表示

RepFlow: Reciprocal Supervision Improves Generation and Representation in Flow Models

Weili Zeng, Feng Tian, Shengqi Liu, Yichao Yan

arXiv 2609.33217首次发表:更新:

AI 中文总结

RepFlow通过互惠学习,从生成器演化状态中提取表示并反馈指导生成,无需外部教师,显著降低FID并提升线性探针准确率,同时支持分布级监督。

AI 中文摘要

生成模型通过去噪学习视觉结构,但其内部状态与噪声水平和网络深度纠缠在一起,使得难以从生成器本身获得稳定的视觉表示。我们提出RepFlow,从生成器的演化计算中学习这种表示,并利用它来指导生成。具体而言,我们训练一个独立的、无时间步的编码器,通过一个时间步条件预测器,从掩蔽的干净图像中恢复跨深度和噪声水平的生成器状态。通过从编码器输入中排除参考噪声和掩蔽内容,我们鼓励编码器提炼出从可见图像上下文可恢复且能预测生成器状态的视觉信息。从生成器演化状态中学习到的表示随后被反馈以指导和改进生成器,其更新后的状态又为进一步的表示学习提供监督,形成一个互惠学习过程。这一互惠过程提升了多步生成和表示质量,通过在ImageNet上使用潜空间SiT和像素空间JiT的冻结线性探针进行评估,无需外部预训练的表示教师。在两种无条件设置中,FID相对于原生训练降低了$19.2$--$40.8\%$,而线性探针准确率相较于最佳搜索的原始生成器特征提高了6.80--10.04个百分点。学习到的表示还作为一步JiT后训练中的分布度量,将其作用从实例级对齐扩展到分布级监督。

英文摘要

Generative models learn visual structure through denoising, yet their internal states are entangled with both noise level and network depth, making it difficult to obtain a stable visual representation from the generator itself. We introduce RepFlow, which learns such a representation from the generator's evolving computation and uses it to guide generation. Specifically, a separate timestep-free encoder is trained, through a timestep-conditioned predictor, to recover generator states across depths and noise levels from masked clean images. By excluding the reference noise and masked-out content from the encoder's input, we encourage the encoder to distill visual information that is recoverable from the visible image context and predictive of generator states. The representation learned from the generator's evolving states is then fed back to guide and improve the generator, whose updated states provide supervision for further representation learning, forming a reciprocal learning process. This reciprocal process improves multi-step generation and representation quality, as measured by frozen linear probing on ImageNet with latent-space SiT and pixel-space JiT, without an externally pretrained representation teacher. Across the two unconditional settings, FID decreases by $19.2$--$40.8\%$ relative to native training, while linear-probe accuracy improves by 6.80--10.04 percentage points over the best searched raw generator features. The learned representation also serves as a distributional metric for one-step JiT post-training, extending its role from instance-level alignment to distribution-level supervision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑