arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

连续对抗平均流迁移

Continuous Adversarial MeanFlow Transfer

Yara Bahram, Zahra Dehghani, Mélodie Desbos, Eric Granger, Pablo Piantanida, Mohammadhadi Shateri

arXiv 2608.19540首次发表:更新:

AI 中文总结

本研究提出MeanFlow-Transfer与Continuous Adversarial MeanFlow,解决有限数据下新域生成器训练的适配与加速问题,在四个源模型适配五个目标域时,FID等指标相当或更优且NFEs最多减125倍,少步FID平均提29%。

AI 中文摘要

在数据有限的新域上训练快速生成器仍面临两个挑战:其一,将预训练的扩散或流模型适配到新域时,未解决其高成本的多步采样问题,且现有加速方法与源参数化(ε、x、v或u)绑定,导致异构预训练模型无通用加速目标;其二,对抗式精炼虽被证明对少步质量有效,但仅针对瞬时速度流,而非平均流(MeanFlow,MF)模型预测的有限区间平均速度。我们解决了这两个问题,提出MeanFlow-Transfer,将异构源输出映射到共享速度表示,用其从源权重初始化MF生成器,并在目标域上优化MF目标,该方法在广泛的预训练模型中,将适配与加速统一到单个训练循环。随后,我们引入Continuous Adversarial MeanFlow(CAMF)作为训练后阶段,将连续对抗流模型从瞬时速度扩展到MF的有限区间平均速度,CAMF对比真实与预测区间端点间学习势的变化,恢复MF回归平均掉的精细细节,且在区间消失极限下退化为瞬时准则。将四个基于ImageNet的源模型——DiT(ε)、SiT(v)、JiT(x)、iMF(u)——适配到五个目标域,带CAMF的MF-T在FID和FDD上与微调教师相当或更优,同时神经函数评估次数(NFEs)最多减少125倍,且CAMF平均提升MF-T少步FID达29%。

英文摘要

Training fast generators on new domains with limited data remains challenging for two reasons. First, adapting a pretrained diffusion or flow model to a new domain leaves its costly multi-step sampling unaddressed, and existing acceleration methods are tied to the source parameterization--$ε$, $x$, $v$, or $u$--leaving heterogeneous pretrained models with no common acceleration target. Second, while adversarial refinement is proven effective for few-step quality, it is formulated only for instantaneous-velocity flows, not for the finite-interval average velocities that MeanFlow (MF) models predict. We address both problems. We propose MeanFlow-Transfer, which maps heterogeneous source outputs into a shared velocity representation, uses it to initialize an MF generator from the source weights, and optimizes an MF objective on the target domain. This unifies adaptation and acceleration in a single training loop across a broad range of pretrained models. We then introduce Continuous Adversarial MeanFlow, a post-training stage that extends continuous adversarial flow models from instantaneous velocities to MF's finite-interval average velocities. CAMF contrasts changes in a learned potential between real and predicted interval endpoints, recovering fine detail that MF regression averages away, and reduces to the instantaneous criterion in the vanishing-interval limit. Adapting four ImageNet-based source models--DiT ($ε$), SiT ($v$), JiT ($x$), iMF ($u$)--to five target domains, MF-T with CAMF matches or exceeds the fine-tuned teacher in FID and FDD at up to $125\times$ fewer Neural Function Evaluations (NFEs), while CAMF improves MF-T's few-step FID by $29\%$ on average.

CommentsPaper under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑