arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

稀疏先验用于高效分布学习

Sparse Priors for Efficient Distribution Learning

Saumya Goyal, Barnabás Póczos

arXiv 2609.20883首次发表:更新:

发表机构

Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对高维分布学习样本复杂度随维度恶化的问题,提出稀疏先验类及稀疏维度度量,证明在k-稀疏先验下可实现O(√(k/n))的贝叶斯风险界,并扩展到学习采样,从而克服维数灾难。

AI 中文摘要

尽管生成式人工智能技术如今被广泛使用并取得成功,但从 $n$ 个样本中学习一个在 $d$ 维空间中支撑的分布的理论保证却退化为 $O(n^{-1/\Theta(d)})$,尽管这被证明是极小化最优的。我们假设现有的界过于悲观,因为平滑性假设不足以捕捉真实应用中经常出现的分布结构。因此,我们引入了稀疏先验类,并定义了“稀疏维度”作为所有分布空间上先验稀疏性的度量。我们表明,在 $k$-稀疏先验下的分布学习,在常见距离度量下实现了 $\Omega(\sqrt{k/n})$ 的贝叶斯风险下界,并在温和的额外假设下,证明了 TV 距离的匹配上界(在 $n,k$ 中渐近地达到对数项)。我们证明了贝叶斯设置中分布学习与学习采样的统计等价性,因此我们的结果也适用于学习采样。虽然 $k$ 仍然可以依赖于维度 $d$ 或内在维度的概念,但我们的结果表明,在适当的先验下学习克服了关于 $n$ 依赖的维数灾难。

英文摘要

Despite the widespread use and success of generative AI techniques today, theoretical guarantees on learning a distribution supported in $d$ dimensions from $n$ samples degrade as $O(n^{-1/Θ(d)})$, though shown to be minimax optimal. We hypothesize that present bounds are too pessimistic because smoothness assumptions are not enough to capture the structure of distributions that often appear in real applications. Consequently, we introduce the class of sparse priors and define the "Sparse Dimension" as a measure of sparsity of a prior over the space of all distributions. We show that distribution learning under a $k$-sparse prior achieves a Bayesian risk lower bound of $Ω(\sqrt{k/n})$ under common distance metrics, and show a matching (up to logarithmic terms asymptotically in $n,k$) upper bound for the TV distance under mild additional assumptions. We show the statistical equivalence of distribution learning and learning to sample in the Bayesian setting so that our results apply to learning to sample as well. While $k$ can still depend on the dimension $d$, or a notion of intrinsic dimension, our results show that learning under an appropriate prior overcomes the curse of dimensionality with respect to the dependence on $n$.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑