arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26860cs.LG

用于视觉生成的摊销矩匹配

Amortized Moment Matching for Visual Generation

Wenze Liu, Xintao Wang, Pengfei Wan, Xiangyu Yue

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出摊销矩匹配方法,构建AMFD损失,用于视觉生成,在ImageNet、GenEval等基准上性能优于FD基线及FLUX.2 4B模型,实现高效高维数据生成。

中文摘要 AI 辅助

我们提出摊销矩匹配方法,利用神经网络学习数据矩作为分布训练信号。通过将扩散去噪器通过多项式投影,建立了矩摊销的通用框架,证明n次投影可明确识别至n+1阶的数据矩。从可处理的仿射情况出发,我们实例化了摊销弗雷歇距离(AMFD)损失。与依赖显式边际矩计算的FD损失不同,AMFD可通过交替、无矩阵优化流程动态学习条件矩,轻松扩展至高维数据。当作用于全局表示特征时,AMFD是强大的训练后目标;经验表明,其神经形式比精确统计匹配的训练动态更稳健,在FDr⁶指标上显著优于FD基线,在ImageNet上实现更优的单步生成。此外,它可在原生生成空间内直接探索,表明前两矩仅能在语义强的空间中识别目标分布。最后,扩展至文本到图像生成时,AMFD的条件感知特性大幅提升了指令遵循能力,使我们的单步模型在GenEval基准上优于多步FLUX.2 [klein] 4B教师模型,同时在PickScore上达到同等性能。代码和检查点可在this https URL获取。

英文摘要

We propose amortized moment matching, utilizing neural networks to learn data moments as distributional training signals. By casting diffusion denoisers through polynomial projections, we establish a general framework for moment amortization, revealing that an $n$-th degree projection explicitly identifies data moments up to order $n+1$. Derived from the tractable affine case, we instantiate the Amortized Fréchet Distance (AMFD) loss. Unlike FD-loss which relies on explicit marginal moment calculations, AMFD is able to dynamically learn conditional moments via an alternating, matrix-free optimization pipeline that effortlessly scales to high-dimensional data. When operating on global representation features, AMFD serves as a powerful post-training objective; empirically, its neural formulation yields more robust training dynamics than exact statistical matching, substantially surpassing the FD baseline on the FDr$^6$ metric and achieving superior one-step generation on ImageNet. Furthermore, it unlocks direct exploration within native generative spaces, suggesting that the first two moments can identify target distributions only in spaces with strong semantics. Finally, when scaled to text-to-image generation, the condition-aware nature of AMFD unlocks massive gains in instruction-following capabilities, enabling our one-step models to outperform their multi-step FLUX.2 [klein] 4B teachers on the GenEval benchmark while achieving on-par performance on PickScore. Code and checkpoints are available at https://github.com/poppuppy/amfd.

发表机构

  • MMLab, CUHK(香港中文大学MMLab)
  • Kling Team, Kuaishou Technology(快手科技影幻团队)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑