所罗门诺夫归纳法与奇异积分
Solomonoff Induction and Singular Integrals
浏览论文内容
中文总结 AI 辅助
该研究将所罗门诺夫分布与奇异积分结合,构造单调图灵机使半测度匹配可计算贝叶斯模型证据,推导得出学习系数作为log n系数出现在所罗门诺夫分布的码长上界中。
中文摘要 AI 辅助
所罗门诺夫分布M通过对所有输出以有限二进制串z开头的程序求和,将长度为ℓ的程序按2^(-ℓ)加权,为有限二进制串z分配先验概率,因此可能的串是那些有许多短解释的串,这似乎是一种纯粹的离散复杂度概念。然而,黎曼和也是可数的。我们证明,定义M的和经适当重组后,包含近似任何可计算统计模型贝叶斯证据的黎曼和。这些证据积分是奇异积分,其渐近性由奇异学习理论支配,并通过奇异学习理论由代数几何的不变量支配。具体而言,对于每个可计算贝叶斯模型,我们构造一个单调图灵机,其诱导的半测度与证据Z_n相差一个统一乘法常数。如果该模型还满足渡部(Watanabe)自由能渐近性的假设,那么从真实分布独立同分布抽取的样本X^n=X_1⋯X_n满足−log M(X^n) ≤ nL_n(w_0)+λlog n−(m−1)log log n+O_P(1),其中L_n是经验损失,w_0是最优参数,λ是学习系数,m是其重数。因此,作为模型简单性几何度量的学习系数,在所罗门诺夫分布中表现为码长上界中log n的系数。
英文摘要
The Solomonoff distribution $M$ assigns an a priori probability to a finite binary string $z$ by summing over all programs whose output begins with $z$, weighting a program of length $\ell$ by $2^{-\ell}$. Thus, likely strings are those with many short explanations. This appears to be a purely discrete notion of complexity. However, Riemann sums are also countable. We show that the sum defining $M$, suitably reorganised, contains Riemann sums approximating the Bayesian evidence of any computable statistical model. These evidence integrals are singular integrals whose asymptotics are governed by Singular Learning Theory and, through it, by invariants of algebraic geometry. Concretely, for every computable Bayesian model, we construct a single monotone Turing machine whose induced semimeasure agrees with the evidence $Z_n$ up to a uniform multiplicative constant. If the model also satisfies the hypotheses of Watanabe's free-energy asymptotics, then a sample $X^n=X_1\cdots X_n$ drawn i.i.d. from the true distribution satisfies $-\log M(X^n) \le nL_n(w_0)+λ\log n-(m-1)\log\log n+O_{\mathbb{P}}(1)$, where $L_n$ is the empirical loss, $w_0$ is an optimal parameter, $λ$ is the learning coefficient, and $m$ is its multiplicity. Thus the learning coefficient, a geometric measure of model simplicity, appears within the Solomonoff distribution as the coefficient of $\log n$ in an upper bound on code length.