发表机构
University of Michigan; Google DeepMind; University of California, Berkeley(密歇根大学; 谷歌DeepMind; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究流匹配模型从有限样本学习未知分布的统计泛化,推导出基于内在维数的Wasserstein距离误差界,证明其适应数据固有几何并缓解维数灾难,为经验成功提供理论解释。
AI 中文摘要
尽管流匹配模型在经验上取得了显著成功,但其统计泛化保证仍不完善。现有分析通常对估计的速度场施加限制性假设,并产生无法反映真实数据(如自然图像和分子几何结构)中常见的内在低维结构的收敛速率。在本工作中,我们研究了流匹配模型从有限样本中学习未知分布$P_{\mathrm{data}}$的统计泛化性质。我们推导了学习到的生成分布在Wasserstein-$p$距离下(对所有$p\geq 1$)的有限样本误差界。具体而言,给定来自$P_{\mathrm{data}}$的$n$个独立同分布样本,我们证明,对于每个$d>d_p^\ast(P_{\mathrm{data}})$以及适当选择的网络架构和超参数,学习到的分布$\widehat{P}^{\mathrm{FM}}$满足$ \mathbb{W}_p(\widehat{P}^{\mathrm{FM}},P_{\mathrm{data}}) \lesssim n^{-1/d}+n^{-1/(2p)}\bigl(\log(1/\xi)\bigr)^{1/(2p)}$,且概率至少为$1-\xi$,其中$d_p^\ast(P_{\mathrm{data}})$表示目标测度的Wasserstein-$p$维数。我们的结果表明,流匹配自然适应数据的固有几何结构,并缓解维数灾难,因为收敛指数取决于内在维数而非环境维数。这些保证在高维场景中仍然有意义,并为流匹配在结构化数据分布上的经验成功提供了理论解释,且其假设比现有分析中的假设宽松得多。
英文摘要
Despite the remarkable empirical success of flow-matching models, their statistical generalization guarantees remain underdeveloped. Existing analyses often impose restrictive assumptions on the estimated velocity field and yield convergence rates that fail to reflect the intrinsic low-dimensional structure common in real data, such as natural images and molecular geometries. In this work, we study the statistical generalization of flow-matching models for learning an unknown distribution $P_{\mathrm{data}}$ from finitely many samples. We derive finite-sample error bounds on the learned generative distribution, measured in the Wasserstein-$p$ distance, for all $p\geq 1$. Specifically, given $n$ i.i.d. samples from $P_{\mathrm{data}}$, we show that, for every $d>d_p^\ast(P_{\mathrm{data}})$ and appropriately chosen network architectures and hyperparameters, the learned distribution $\widehat{P}^{\mathrm{FM}}$ satisfies $ \mathbb{W}_p(\widehat{P}^{\mathrm{FM}},P_{\mathrm{data}}) \lesssim n^{-1/d}+n^{-1/(2p)}\bigl(\log(1/ξ)\bigr)^{1/(2p)}$ with probability at least $1-ξ$, where $d_p^\ast(P_{\mathrm{data}})$ denotes the Wasserstein-$p$ dimension of the target measure. Our results demonstrate that flow matching naturally adapts to the intrinsic geometry of data and mitigates the curse of dimensionality, as the convergence exponent depends on the intrinsic rather than ambient dimension. These guarantees remain meaningful in high-dimensional regimes and provide a theoretical explanation for the empirical success of flow matching on structured data distributions under substantially more relaxed assumptions than those in existing analyses.