发表机构
University of Auckland; Strong AI Lab(奥克兰大学; 强人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出分层所罗门诺夫归纳法(HSI),将德菲涅蒂定理应用于所罗门诺夫归纳法(SolInd),证明HSI等价于SolInd,其超额误差受生成器复杂度约束,随数据集增长收敛至0,是适用于给定数据集的无界序列预测理想模型。
AI 中文摘要
所罗门诺夫归纳法(SolInd)提供了一种理想的无界先验序列预测模型,但无法自然描述大型语言模型所执行的、从给定训练数据集出发的外推过程。我们将德菲涅蒂的可交换分布定理应用于SolInd,得到了我们所称的分层所罗门诺夫归纳法(HSI),该模型在所有所罗门诺夫先验上维护一个超先验,可基于先前观测到的序列进行条件调整。我们扩展了Wood等人的证明——半测度的通用混合等价于SolInd——以表明这些混合的通用混合也等价,从而证明HSI=SolInd。我们还证明,HSI在任意分布上相对于其真实生成器的超额误差,受限于该生成器在超先验中的复杂度。该结果可直接与SolInd的预测误差受限于被预测序列的柯尔莫哥洛夫复杂度相比较,且当数据集规模增长时,HSI的平均超额误差会收敛至0,最终实现最优预测。我们声称,HSI是一种理想的无界序列预测模型,适用于给定数据集的场景,正如SolInd是针对单个序列的理想模型一样。
英文摘要
Solomonoff Induction, or SolInd, provides an ideal unbounded model of a priori sequence prediction but cannot naturally describe extrapolation from a given training dataset, as performed by Large Language Models. We apply de Finetti's theorem on exchangeable distributions to SolInd to produce what we call Hierarchical Solomonoff Induction, or HSI, which maintains a hyperprior over all Solomonoff priors that can be conditioned on previously observed sequences. We extend Wood et al.'s proof that universal mixtures of semimeasures are equivalent to SolInd to show that universal mixtures of these mixtures are also equivalent, proving that HSI=SolInd. We also prove that HSI's excess error on any distribution, compared to its true generator, is bounded by that generator's complexity in the hyperprior. This result is directly comparable to SolInd's prediction error being bounded by the Kolmogorov complexity of the sequence being predicted, and forces HSI's average excess error to converge to 0 as a dataset grows, leading to optimal prediction in the limit. We claim that HSI is an ideal unbounded model of sequence prediction given a dataset in the same way that SolInd is ideal over individual sequences.