发表机构
The University of Texas at Austin; Cornell University(德克萨斯大学奥斯汀分校; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究得分-熵离散扩散(SEDD)的统计极限,针对均匀和掩码离散扩散模型建立极小极大下界,提出匹配下界的MLE阈值估计器,证明SEDD经合适设置可达到接近最优的极小极大样本复杂度。
AI 中文摘要
离散扩散模型在自然语言数据、图结构数据等多种数据集上展现出优异性能,其中得分-熵离散扩散(SEDD)变体的实验结果尤为突出。SEDD通过迭代评估一系列具体得分函数生成新样本,这些得分函数通过最小化得分-熵损失学习得到。现有离散扩散的理论文献多聚焦于小得分估计误差假设下SEDD的采样效率,近期研究开始探究得分估计本身的有限样本性质。本文另辟蹊径,研究具体得分估计的基础统计极限,聚焦两种应用最广泛的离散扩散模型:均匀离散扩散和掩码离散扩散。我们在得分-熵损失下建立极小极大下界,提出一种基于极大似然估计(MLE)的阈值估计器,其性能与该下界匹配度可达依赖邻域密度比的常数和多项式对数因子。进一步证明,对于任意目标分布,该密度比在均匀和掩码离散扩散模型下均受自然控制,使聚合得分估计误差的极小极大上下界几乎匹配。结果表明,通过合适的初始化和离散化,SEDD可达到接近最优的极小极大样本复杂度,该复杂度以目标分布与生成分布间的KL散度衡量。
英文摘要
Discrete diffusion models have demonstrated strong performance across a range of datasets, including natural language data and graph-structured data. Among many variants, score-entropy discrete diffusion (SEDD) has achieved particularly strong empirical results. In SEDD, new samples are generated by iteratively evaluating a sequence of concrete score functions, which are learned by minimizing a score-entropy loss. While much of the prior theoretical literature on discrete diffusion has focused on the sampling efficiency of SEDD under the assumption of small score estimation error, recent work has begun to investigate the finite-sample properties of score estimation itself. In this work, we take a different route by investigating the fundamental statistical limits of concrete score estimation. We focus on uniform and masking discrete diffusions, two of the most widely adopted discrete diffusion models. We establish a minimax lower bound under the score-entropy loss, and propose an MLE-based thresholding estimator that matches this lower bound up to constant and polylogarithmic factors that depend on neighboring density ratios. We further show that, for any target distribution, this density ratio is naturally controlled under both uniform and masking discrete diffusion models, yielding nearly matching minimax lower and upper bounds for the aggregated score estimation error. Our results imply that, with appropriate initialization and discretization, SEDD can achieve nearly optimal minimax sample complexity, as measured by the KL divergence between the target and generated distributions.
Comments26 pages, 3 figures