arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14608math.STcond-mat.stat-mechmath-phmath.MPphysics.data-anstat.MEstat.TH

采样的玻尔兹曼结构:固有p值及其涌现的闭式表达式

The Boltzmann structure of sampling: Intrinsic $p$-value and its emergent closed-form expression

Orestis Loukas

AI总结:

该研究针对受测量分辨率限制于有限类别的可观测变量,通过前向采样的组合构造定义固有离散p值,利用鞍点技术推导其连续极限近似,结合信息几何与球对称性,得到可通过χ²_k分布高效计算的半闭式p值统计量。

AI中文摘要:

我们考虑可观测变量X,其在采样数据集中的实现受限于有限个可区分类别,例如受测量分辨率限制,这些类别位于其可能无限的理论域内。给定可观测类别的概率,我们研究通过前向采样过程生成的各类可能数据集的粗粒度概率质量。这种组合构造诱导出一种固有离散p值,它直接由采样过程定义,而非通过可观测变量上的额外概率结构。指定d+k个任意函数g_α(X)的样本均值,定义了一个线性数据集族,其概率质量为关联多面体内整数格点上的加权和。为克服大N下难以处理的组合问题,我们通过鞍点技术在多项普适类的前向采样连续极限中推导出一种密度,用于近似这些概率质量。作为示例,我们考虑条件采样,其中d个结构均值固定,k个均值在容许数据集上变化。由鞍点密度涌现出的信息几何,结合固有p值构造在大N下产生的球对称性,使我们能通过χ²_k分布在拉普拉斯近似中高效计算p值。所得统计量为2N乘以对应结构线性族与观测线性族相关信息投影之间的Kullback-Leibler散度的半解析表达式,这些投影可通过标准数值例程高效计算,只要样本均值表现足够良好。

英文摘要:

We consider observables $X$ whose realizations in a sampled dataset are restricted, for example by measurement resolution, to a finite set of distinguishable categories within their possibly infinite theoretical domain. Given the probabilities of observable categories, we study the coarse-grained probability mass of families of possible datasets generated through a forward sampling process. The combinatorial construction induces an intrinsically discrete $p$-value defined directly from the sampling process rather than through additional probabilistic structure on observables. Specifying the sample means of $d+k$ arbitrary functions $g_α(X)$ defines a linear family of datasets whose probability mass is obtained as a weighted sum over integer lattice points contained within the associated polyhedron. To overcome the intractable large-$N$ combinatorics, we derive via saddle-point techniques a density approximating these probability masses in the continuum limit of forward sampling within the multinomial universality class. As a demonstration, we consider conditional sampling, where $d$ structural means are fixed while $k$ means vary over admissible datasets. The information geometry emerging from the saddle-point density, together with the spherical symmetry arising at large $N$ from the intrinsic $p$-value construction, enables efficient computation of the $p$-value in the Laplace approximation via the $χ^2_k$ distribution. The resulting statistic is given by the semi-analytic expression $2N$ times the Kullback-Leibler divergence between the information projections associated with the corresponding structural and observed linear families. These projections can be computed efficiently via standard numerical routines converging for sufficiently well-behaved sample means.

↑