探索式建模:解锁第三大预训练轴与端到端生成
Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- UIUC(伊利诺伊大学厄巴纳-香槟分校)
- Harvard(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出探索式建模(XMs),为生成式模型增添第三预训练轴,可提升多域性能与效率,还能实现端到端重构生成建模,推理步骤大幅减少。
AI中文摘要:
由AlexNet开启的深度学习革命让我们认识到,端到端训练优于将问题分解为手工设计的阶段。然而生成式建模仍是例外——尽管生成式模型能力极强,它们仍未实现端到端训练。这是因为生成式建模的核心是处理多模态分布,现有可扩展方法均通过分解生成过程来处理,这阻碍了端到端生成。本研究中,我们提出探索式建模(Explorative Modeling)这一新范式,它转而分解训练循环,探索模型生成与数据间的K个候选匹配,并基于最优匹配进行训练,使预测聚焦于模态而非模糊它们。我们发现探索式模型(Explorative Models, XMs)在两种场景下表现出色:其一,增加探索度为现有生成式模型增添了参数与数据之外的第三大预训练轴,且探索度的扩展在连续域与离散域(图像、视频、语言)均能持续提升性能。值得注意的是,探索带来的增益随规模增长,当数据规模扩大时,增益从7%升至36%;当模型规模增长时,增益从13%升至23%,计算量增至3倍时,效率增益翻倍有余。具体而言,探索将FLOP效率提升4.1倍、样本效率提升6.2倍、参数效率提升47%,在无引导的情况下,将最强的图像生成方案在ImageNet上的FID提升至接近当前最优的1.43,实现了现有模型端到端程度的扩展,并解锁了泛化能力的扩展。其二,XMs支持端到端重构生成式建模,在控制任务上与扩散模型表现相当,但推理步骤减少16至256倍。综上,这些结果确立了XMs既是现有生成式模型的新预训练轴,也是独立的端到端生成式建模范式。
英文摘要:
The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages. Generative modeling, however, has remained the exception-despite generative models being remarkably capable, they are still not trained end-to-end. This is because, at its core, generative modeling is about handling distributions with many modes, and existing scalable approaches handle this the same way, by factoring the generation procedure, which prevents end-to-end generation. In this work, we introduce Explorative Modeling, a new paradigm that instead factors the training loop, exploring K candidate matches between model generations and data, and training on the best, so predictions commit to modes rather than blurring them. We find Explorative Models (XMs) useful in two settings. First, increasing exploration adds a third pretraining axis beyond parameters and data for existing generative models-where scaling exploration monotonically improves performance across both continuous and discrete domains (images, video, and language). Notably, gains from exploration increase with scale, climbing from 7% to 36% as data scales and from 13% to 23% as models grow, with efficiency gains more than doubling at 3x the compute. Concretely, exploration improves FLOP efficiency by 4.1x, sample efficiency by 6.2x, parameter efficiency by 47%, lifts the strongest of image-generation recipes to a near-state-of-the-art 1.43 FID on ImageNet without guidance, enables scaling how end-to-end existing models are, and unlocks scaling generalization. Second, XMs enable end-to-end reconstructive generative modeling, matching diffusion on control tasks with 16-256x fewer inference steps. Together, these results establish XMs as both a new pretraining axis for existing generative models and a standalone end-to-end generative modeling paradigm.