arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19332cs.LGcs.CV

ROMS-IMLE:一种用于竞争性单步生成建模的极简方法

ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling

Chirag Vashist, Ke Li

首次发表
浏览论文内容

中文总结 AI 辅助

质疑生成模型中逐步变换噪声分布的必要性,采用极简方法,选IMLE为训练目标、卷积网络为模型,经添加关键要素得到单步参数高效生成模型,在ImageNet 256上取得良好效果。

中文摘要 AI 辅助

生成模型历经多代演变,从VAE/GAN到扩散/流匹配,基础技术愈发复杂。基于扩散模型和流匹配的成功,人们普遍认为需经多次小变换将噪声分布逐步转换为数据分布。本文对此提出质疑,采用极简方法设计竞争性生成模型。从最基本要素出发,选择隐式最大似然估计(IMLE)作为训练目标,摒弃变分推理、对抗训练等复杂方法;选用适度规模的卷积网络而非Transformer。经审慎添加关键要素,得到单步参数高效的生成模型,在ImageNet 256上FID达2.56,同时具备良好的精度和召回率。

英文摘要

Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching. Along the way, the underlying techniques have become more complicated and various beliefs about what drives strong empirical performance have taken hold. Due to the success of diffusion models and flow matching, one of the more common beliefs is the importance of transforming the noise distribution to the data distribution gradually through many small transformations. We ask whether this is truly necessary, and take a minimalist approach to designing a competitive generative model. We start with the bare-bones essentials, namely just a training objective and a model. We purposefully make both simple. For the training objective, we choose Implicit Maximum Likelihood Estimation (IMLE), and eschew more complicated alternatives such as variational inference, adversarial training and numerical integration. For the model, we eschew transformers and instead choose a moderately sized convolutional network. Then we judiciously added elements that are truly essential, which surprisingly do not include iterative denoising. The result is a single-step parameter-efficient generative model that produces high quality samples at fast speed: it achieves an FID of 2.56 on ImageNet 256 and simultaneously attains good precision and recall.

↑