发表机构
University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对离散扩散模型统计界随环境空间大小退化的问题,本文证明均匀离散扩散的TV误差率由有效支持大小决定且极小极大最优,KL率仅多对数因子,从而避免维数灾难。
AI 中文摘要
离散扩散模型已成为离散乘积空间上生成建模的一个实际成功的框架,但其统计泛化性质仍鲜为人知。诸如文本或生物序列等离散现实世界数据,由于语义或物理约束,往往集中在极其庞大的环境空间的一小部分上,但现有界未能捕捉这种分布结构,而是随环境空间大小缩放,导致几乎无意义的误差界。我们针对均匀离散扩散(与掩码扩散并列的两大离散扩散范式之一)填补了这一空白,通过推导由有效支持大小$s_n(P_0)$(一种依赖于样本量的分布复杂度度量)所决定的统计保证。给定来自$[K]^d$上未知数据分布$P_0$的$n$个独立同分布(i.i.d.)样本,我们证明,在适当选择网络大小和超参数的情况下,期望总变差(TV)损失按$O(\sqrt{s_n(P_0)/n})$缩放,而期望Kullback--Leibler(KL)散度以$O(\frac{1}{n}s_n(P_0)\log(eK^d/s_n(P_0))\log n)$为界。此外,我们证明TV率是极小极大最优的,KL率在$\log n$因子内是极小极大最优的。综合这些上界和下界,表明均匀离散扩散成功避免了小有效支持大小分布的维数灾难:TV误差率仅通过$s_n(P_0)$依赖于环境状态空间大小,而相应的KL率仅对环境状态空间大小产生额外的对数依赖。
英文摘要
Discrete diffusion models have emerged as a practically successful framework for generative modeling on discrete product spaces, yet their statistical generalization properties remain poorly understood. Discrete real-world data such as text or biological sequences often concentrate on a small fraction of the astronomically large ambient space because of semantic or physical constraints, but existing bounds fail to capture this distributional structure and instead scale with the size of the ambient space, giving rise to almost vacuous error bounds. We address this gap for uniform discrete diffusion, one of the two dominant discrete diffusion paradigms alongside masking diffusion, by deriving statistical guarantees governed by the effective support size $s_n(P_0)$, a sample-size-dependent measure of distributional complexity. Given $n$ independent and identically distributed (i.i.d.) samples from an unknown data distribution $P_0$ on $[K]^d$, we show that, with appropriate choices of network size and hyperparameters, the expected total variation (TV) loss scales as $O(\sqrt{s_n(P_0)/n})$, while the expected Kullback--Leibler (KL) divergence is bounded by $O(\frac{1}{n}s_n(P_0)\log(eK^d/s_n(P_0))\log n)$. Furthermore, we show that the TV rate is minimax optimal and that the KL rate is minimax optimal up to a factor of $\log n$. Together, these upper and lower bounds show that uniform discrete diffusion successfully avoids the curse of dimensionality for distributions with small effective support size: the TV error rate depends on the ambient state-space size only through $s_n(P_0)$, while the corresponding KL rate incurs only an additional logarithmic dependence on the ambient state-space size.