发表机构
Booth School of Business, University of Chicago(芝加哥大学布斯商学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对局部依赖的高维类别分布采样,提出基于离散扩散的端到端学习方法,通过钉扎分解和权重共享得分网络获得显式样本复杂度保证,并在实验中优于全连接网络。
AI 中文摘要
统计学、经济学和物理学中的许多应用需要从具有局部依赖结构的高维类别分布中进行采样。例如,有限记忆语言模型、统计物理中的伊辛(Ising)和波茨(Potts)系统以及蛋白质折叠等。在现代机器学习中,离散扩散已成为一种灵活且具有强大实证性能的采样方法。受此启发,我们开发了具有端到端样本复杂度界的离散扩散学习方法,该方法在局部依赖下采用均匀噪声化,并通过低阶马尔可夫随机场(MRFs)进行建模。我们的主要技术洞见是离散得分的一种新的“钉扎分解”(pinning decomposition)。它表明,与连续扩散不同,得分可分解为多个分量,其中对时间的依赖与对目标的依赖以乘法方式分离。基于此分解,我们提出了一种“权重共享神经得分学习器”(weight-sharing neural score learner),并将其与τ-跳跃(τ-leaping)相结合,以获得端到端的采样过程。与现有采样分析中常见的将得分学习误差视为黑箱输入不同,我们从有限数据出发研究得分学习误差,并推导出具有显式依赖词汇表大小、马尔可夫随机场交互阶数和样本量的最优采样保证。此外,我们的策略在均匀噪声水平上训练单个得分网络,而将采样离散化留到推理时选择。这使得同一训练模型能够在推理时间预算变化时,以精度换取计算成本。在波茨(Potts)、伊辛(Ising)和树结构模型上的数值实验表明,权重共享得分网络在采样长序列方面优于全连接网络。
英文摘要
Many applications in statistics, economics, and physics require sampling from high-dimensional categorical distributions with local dependence structures. Examples include finite memory language models, Ising and Potts systems in statistical physics and protein folding, etc. In modern machine learning, discrete diffusions have emerged as a flexible approach for sampling such data, with strong empirical performance. Motivated by this, we develop learning methods with end-to-end sample complexity bounds for discrete diffusion with uniform noising under local dependence, which we model through low order Markov random fields (MRFs). Our main technical insight is a new \emph{pinning decomposition} of the discrete score. It shows that unlike in continuous diffusions, the score decomposes into components where the dependence on time separates multiplicatively from the dependence on the target. Building on this decomposition, we propose a \emph{weight-sharing neural score learner} and combine it with $τ$-leaping to obtain an end-to-end sampling procedure. Rather than treating score-learning error as a black-box input, as is common in existing sampling analyses, we study the score learning error from finite data and derive optimal sampling guarantees with explicit dependence on the vocabulary size, the interaction order of the MRF, and the sample size. Moreover, our strategy trains a single score network across uniform noise levels while leaving the sampling discretization to be chosen at inference-time. This allows the same trained model to trade accuracy for computational cost as inference-time budgets vary. Numerical experiments on Potts, Ising, and tree-structured models show that weight-sharing score networks outperform fully connected ones for sampling long sequences.
Comments83 Pages, 3 Figures, 4 Tables