Barron最优传输I:生成建模
Barron Optimal Transport I: Generative Modeling
浏览论文内容
中文总结 AI 辅助
该研究引入Barron最优传输框架,将其用于生成建模,量化了扩散生成建模在Barron几何中的次优性,还将该几何扩展至采样应用以优化Stein变分梯度方法。
中文摘要 AI 辅助
受生成建模与采样领域近期应用的启发,我们引入了一种最优测度传输框架,其中代价函数捕捉神经网络复杂度的概念。在基于传输的生成模型中,参考分布(如高斯分布)的样本会沿常微分方程或随机微分方程映射为目标分布的样本;当时间离散化时,这些方程会被实现为深度残差网络,每个隐藏层近似相关的瞬时速度。因此,给定一对目标测度与参考测度,一个自然的问题是寻找实现该传输的最高效神经表示。我们的研究起点是Benamou和Brenier提出的最优传输(OT)动力学公式,我们将平均动力学L²能量替换为Barron能量[bach2017breaking, ma2022barron]——这是一种衡量用神经隐藏层表示给定向量场复杂度的自然范数,且能捕捉特征学习的自适应特性。这在概率测度空间上定义了一种度量,补充了现有的Wasserstein几何与Stein几何。本研究中,我们在生成建模背景下考察该度量的性质:作为首个应用,我们通过建立高斯分布经神经网络推前生成的数据的超多项式得分近似下界,量化了扩散生成建模在Barron几何中的次优性;随后,我们研究了自适应在替代生成模型研究中的作用。在配套论文[companionpaper]中,我们将Barron传输几何用于采样应用,通过特征自适应扩展Stein变分梯度方法的适用范围。
英文摘要
Motivated by recent applications in generative modeling and sampling, we introduce a framework for optimal measure transport where cost captures the notion of neural network complexity. In transport-based generative models, samples from a reference distribution (e.g. Gaussian) are mapped to samples of a target distribution along ordinary or stochastic differential equations. These are implemented as deep residual networks when discretized in time, where each hidden layer approximates the associated instantaneous velocity. Thus, given a pair of target and reference measures, a natural question is to search for the most efficient neural representation that implements this transport. Our starting point is the kinetic formulation of OT, due to Benamou and Brenier. We replace the average kinetic $L^2$ energy by the \emph{Barron} energy \cite{bach2017breaking, ma2022barron}, a natural norm which measures the complexity of representing a given vector field with a neural hidden layer, and which captures the adaptive properties of feature learning. This defines a metric on the space of probability measures, complementing existing Wasserstein and Stein geometries. In this work we examine the properties of this metric in the context of generative modeling. As a first application, we quantify the suboptimality of diffusion generative modeling in the Barron geometry by establishing super-polynomial score approximation lower bounds for data generated by neural network pushforwards of the Gaussian. We then investigate the benefit of adaptivity as a way to study alternative generative models. In a companion paper \cite{companionpaper} we leverage the Barron transport geometry for sampling applications, extending the scope of Stein variational gradient methods via feature adaptation.