掩码离散扩散中tau-leaping的调度优化
Schedule optimization for tau-leaping in masked discrete diffusion
浏览论文内容
中文总结 AI 辅助
本文分析掩码离散扩散模型中tau-leaping采样器的因子分解误差,通过依赖密度刻画调度优化问题,证明在非退化情形下调度仅改善常数项,在退化情形下可改善渐近阶。
中文摘要 AI 辅助
掩码离散扩散模型通常使用所谓的tau-leaping离散化方法加速,该方法在每个采样步骤并行揭示多个坐标。采样器将每个揭示块的联合条件分布替换为乘积分布,即使预测器完美学习,也会产生因子分解误差$\varepsilon_\text{fact}$。我们分析了在$N$个坐标上具有$K$个采样步骤的标准采样器,其随机块大小依赖于去噪调度。我们的分析使用了$\varepsilon_\text{fact}$的精确积分表示,该表示基于分布相关的依赖密度$\rho$,它记录了随着揭示坐标比例增加,条件依赖如何演变。我们为这一轮廓开发了估计器,并量化了估计误差如何影响调度选择。我们推导了有限$K$优化问题的递归平稳方程,并在单调性条件下刻画了其唯一优化器。在联合极限$N,K\to\infty$下,我们获得了最优极限平滑调度的显式刻画,并量化了相对于确定性规划器的随机块大小的成本。当$\rho_N$一致收敛到严格正的连续轮廓时,在固定平滑调度上优化可以改善领先常数,但不能改善$\varepsilon_\text{fact}$的$N/K$缩放。相反,如果$\rho_N$退化,合适的调度可以相对于均匀调度改善渐近阶。基于平稳过程和可交换混合物的例子说明了这两种机制。
英文摘要
Masked diffusions are popular generative models for discrete distributions. Unlike standard autoregressive sampling, they reveal several coordinates in parallel, approximating each block's joint conditional law by a product of one-coordinate conditionals. The resulting procedure, usually called tau-leaping, reduces computational cost but introduces a factorization error ($\varepsilon_\text{fact}$), even with perfectly learned predictors. We study the resulting tradeoff between generative accuracy and computational cost, focusing on how to choose a denoising schedule to minimize $\varepsilon_\text{fact}$ for a fixed sampling budget. To do so, we establish an exact integral representation of $\varepsilon_\text{fact}$ separating the schedule from the target's dependence structure, summarized by a dependence density $ρ$. This representation yields recursive stationarity equations for optimal schedules and allows us to quantify how estimation errors in $ρ$ affect schedule selection. As the dimension $N$ and sampling budget grow, we characterize the optimal schedule and quantify the cost of random block sizes relative to a deterministic planner. We highlight a fundamental dichotomy: if $ρ$ converges uniformly to a strictly positive continuous profile as $N\to\infty$, schedule optimization can only improve the leading constant of $\varepsilon_\text{fact}$, while if $ρ$ degenerates, schedule optimization can improve the asymptotic order. Examples based on stationary processes and exchangeable mixtures illustrate these regimes.
发表机构
- Bocconi University(博科尼大学)
机构由 AI 辅助整理,请以论文原文为准。