发表机构
University of Pennsylvania; the Wharton School; Department of Statistics and Data Science(宾夕法尼亚大学; 沃顿商学院; 统计与数据科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对均匀与重掩蔽离散扩散模型,提出带留一法去噪器的一阶自适应采样器,证明其采样复杂度由目标分布固有依赖结构决定,而非环境维度,实验验证了维度自适应行为。
AI 中文摘要
离散扩散模型通过实现并行更新,为自回归生成提供了一种有前景的替代方案,但其采样效率很大程度上取决于前向过程和采样器的选择。对于均匀前向过程,现有标准τ跳步采样器的下界与环境维度d呈线性比例关系,这引发了一个问题:这种依赖是否是前向过程固有的。我们否定了这个问题。我们考虑一种基于留一法去噪器的一阶采样器,适用于均匀和重掩蔽过程,其坐标更新可并行执行。在两种情况下,该采样器都能在采样过程中纠正去噪错误,这在同时更新多个坐标时变得必要。我们的主要结果确立了自适应采样保证:忽略对数因子,N = O(DTC(X₀)/ε)个离散化步骤足以实现采样误差O(ε_score + ε),其中ε_score是评分估计的误差。因此,采样复杂度由目标分布的固有依赖结构决定,该结构通过其对偶总相关DTC(X₀)衡量,而非直接由环境维度d决定。我们的分析通过贝叶斯最优辅助采样器进行,该采样器将离散化误差与评分估计误差分离。我们还根据前向过程不同时间不同坐标之间的互信息,推导了离散化误差的精确信息论表示,该表示适用于一般前向过程,在均匀和重掩蔽情况下可由DTC(X₀)控制。对结构化合成分布的数值实验说明了预测的维度自适应行为。
英文摘要
Discrete diffusion models offer a promising alternative to autoregressive generation by enabling parallel updates, but their sampling efficiency can depend strongly on the choice of the forward process and the sampler. For the uniform forward process, existing lower bounds for the standard $τ$-leaping sampler scale linearly with the ambient dimension $d$, raising the question of whether this dependence is intrinsic to the forward process. We answer this question in the negative. We consider a first-order sampler based on the leave-one-out denoiser for uniform and remasking processes whose coordinate updates can be performed in parallel. In both cases, the sampler can correct denoising mistakes during the sampling process, which becomes necessary when many coordinates are updated together. Our main result establishes an adaptive sampling guarantee: up to logarithmic factors, $N = O(\mathrm{DTC}(X_0) / \varepsilon)$ discretization steps suffice to achieve sampling error $O(\varepsilon_{\mathrm{score}}+\varepsilon)$, where $\varepsilon_{\mathrm{score}}$ is the error in score estimation. Thus, the sampling complexity is governed by the intrinsic dependence structure of the target distribution, as measured by its dual total correlation $\mathrm{DTC}(X_0)$, rather than directly by the ambient dimension $d$. Our analysis proceeds through a Bayes-optimal auxiliary sampler that separates discretization error from score-estimation error. We also derive an exact information-theoretic representation of the discretization error in terms of the mutual information between different coordinates of the forward process at different times. This representation applies to general forward processes and, in the uniform and remasking cases, can be controlled by $\mathrm{DTC}(X_0)$. Numerical experiments on structured synthetic distributions illustrate the predicted dimension-adaptive behavior.
Comments37 pages, 4 figures