arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

流匹配中的速度缩放

Velocity Scaling in Flow Matching

Youssef Saied, François Fleuret

arXiv 2610.10823首次发表:更新:

发表机构

University of Geneva; Meta FAIR(日内瓦大学; Meta FAIR)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对流匹配中速度缩放的作用,证明MSE训练不会导致速度幅值不足,提出速度缩放可降低总体时间滞后,据此选择增益使ImageNet-256的FID从28.0降至12.2。

AI 中文摘要

近期研究表明,用增益γ(t)缩放学习到的流匹配速度场v_θ可大幅提升生成质量。此前研究认为,用均方误差(MSE)训练的速度场会系统性低估速度幅值,缩放可修正该误差。本文证明MSE训练不会造成速度幅值不足,反而发现速度缩放会降低总体时间滞后:在模型时间t采样的状态更接近更早时间的训练状态。速度缩放与回退模型时间是解决该总体时间滞后的两种方式。在多种架构和模型规模下,通过测量总体时间滞后并据此选择增益,可显著提升生成质量,在无引导的NFE 25下,ImageNet-256的FID从28.0降至12.2(由相邻增益下FID测量值线性插值估算)。

英文摘要

Scaling a learned flow-matching velocity field $v_θ$ by a gain $γ(t)$ was recently shown to greatly improve generation quality. Prior work argued that velocity fields trained with mean-squared error (MSE) systematically underestimate velocity magnitude and that scaling corrects this error. We show that MSE training does not create a velocity-magnitude deficit. We find instead that velocity scaling reduces population time lag: sampled states at model time $t$ resemble training states from an earlier time. Velocity scaling and moving model time back are two ways to address this population time lag. Across architectures and model sizes, measuring population time lag and using it to select a gain greatly improves generation quality, reducing FID from 28.0 to 12.2 (estimated by linear interpolation between FID measurements at neighboring gains) on ImageNet-256 at NFE 25 without guidance.

Comments41 pages, including appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑