arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14503cs.CCcs.DScs.ITmath.ITmath.PR

随机采样的空间熵下界

Space-Entropy Lower Bounds for Random Sampling

Thomas L. Draper, Feras A. Saad

首次发表
浏览论文内容

中文总结 AI 辅助

研究随机采样算法接近信息理论最优熵界所需的空间,通过图论分析证明对于任意离散目标分布精确且满足特定输入比特消耗条件的算法,需\(\Omega(\log(1/\varepsilon))\)比特空间,还给出多种分布下的具体空间下界。

中文摘要 AI 辅助

我们使用独立同分布均匀比特的熵源证明了精确随机采样的基本空间下界。信息论的一个经典结果表明,生成\(n\)个离散随机变量\(X_1,\dots,X_n\)平均至少需要\(H(X_1,\dots,X_n)\)个输入随机比特,其中\(H\)是香农熵函数。为了接近这个信息理论上最优的熵界,随机采样算法必须使用多少空间呢?我们证明,对于任意离散目标分布精确且每个输出过程期望消耗至多\(H(X_1,\ldots,X_n)+\varepsilon n+o(n)\)个输入比特的任何随机采样算法,必须使用\(\Omega(\log(1/\varepsilon))\)比特的空间。实际上,从单个\(\mathrm{Bernoulli}(1/3)\)分布进行独立同分布采样就已经迫使至少\((1/{5.116201}-o(1))\log(1/\varepsilon)\)比特的空间。如果采样器处理无限多个伯努利分布族,我们展示了至少\(\log(1/\varepsilon)\)比特空间的更紧界。我们还证明了一般独立同分布采样的下界:对于几乎每个\(k\)个结果上的分布,空间至少为\((1/(k + 1)-o(1))\log(1/\varepsilon)\)比特。证明技术基于对任何算法在其状态中可存储信息量的图论分析。有限状态空间迫使状态转移图周围出现短循环,并且围绕此类循环的损失归结为对数概率整数组合的分数部分的丢番图下界。据我们所知,这些结果构成了关于熵高效随机采样的首个已知空间下界。

英文摘要

We prove fundamental space lower bounds for exact random sampling using an entropy source of i.i.d. uniform bits. A classic result from information theory shows that generating $n$ discrete random variables $X_1, \dots, X_n$ requires at least $H(X_1, \dots, X_n)$ input random bits on average, where $H$ is the Shannon entropy function. How much space must a random sampling algorithm use in order to approach this information-theoretically optimal entropy bound? We prove that any random sampling algorithm that is exact for arbitrary discrete target distributions and consumes at most $H(X_1,\ldots,X_n)+\varepsilon n+o(n)$ input bits in expectation for every output process must use $Ω(\log(1/\varepsilon))$ bits of space. In fact, i.i.d. sampling from the single distribution $\mathrm{Bernoulli}(1/3)$ already forces at least $(1/{5.116201}-o(1))\log(1/\varepsilon)$ bits of space. If the sampler handles a family of infinitely many Bernoulli distributions, we show a sharper bound of at least $\log(1/\varepsilon)$ bits of space. We also prove lower bounds for general i.i.d. sampling: for almost every distribution on $k$ outcomes, the space is at least $(1/(k+1)-o(1))\log(1/\varepsilon)$ bits. The proof technique is based on a graph-theoretic analysis of the amount of information that any algorithm can store in its state. Finite state spaces force short cycles around the state-transition graph, and the loss around such cycles reduces to Diophantine lower bounds on fractional parts of integer combinations of log-probabilities. To the best of our knowledge, these results comprise the first known space lower bounds for entropy-efficient random sampling.

↑