AI 中文总结
本文分析流匹配中不同预测目标性能差异的根源,提出跨时间与协方差方向自适应的谱混合参数化,证明其对高斯数据最优,并在多种架构下显著加速优化且几乎不增加训练成本。
AI 中文摘要
流匹配和扩散模型可以被训练来预测不同的量,最常见的是数据$x_1$、源噪声$x_0$或速度$v$。尽管理论上等价,但这些选择可能导致截然不同的性能表现。我们识别出导致这些差异的两个主要驱动因素:源-数据信噪比,以及神经架构引起的信息瓶颈。我们表明,在内在数据维度之外,影响最优参数化最显著的因素是每个数据协方差方向上的特定信噪比。基于这一分析,我们引入了新的\emph{谱混合}参数化方法,该方法能跨时间和数据协方差方向进行自适应;我们证明这些方法对于高斯数据是最优的。我们还表明,架构引起的压缩会改变哪种参数化更易于学习,其中$v$-预测对丢弃方向的敏感度高于$x_1$-预测。跨架构和源尺度的实验表明,我们的谱参数化在不同机制下具有鲁棒性,能显著加速优化,同时与标准参数化相比几乎不增加额外训练成本。
英文摘要
Flow matching and diffusion can be trained to predict different quantities, most commonly the data $x_1$, the source noise $x_0$, or the velocity~$v$. Although theoretically equivalent, these can lead to substantially different performances. We identify two main drivers for these differences: the source--data signal-to-noise ratio, and the information bottleneck induced by the neural architecture. We show that, beyond intrinsic data dimension, the factor affecting the optimal parametrization the most is a certain signal-to-noise ratio in each data covariance direction. From this analysis, we introduce new \emph{spectral hybrid} parameterizations that adapt across time and data covariance directions; we show that these are optimal for Gaussian data. We also show that architecture-induced compression changes which parameterization is easier to learn, with $v$-prediction being more sensitive to discarded directions than $x_1$-prediction. Experiments across architectures and source scales show that our spectral parameterizations are robust across regimes, can substantially accelerate optimization, while incurring essentially no additional training cost compared with standard parameterizations.