发表机构
RIKEN(理化学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示大型语言模型贝叶斯方法因依赖证据下界而缺乏重参数化不变性,导致权重差异显著且方向反转,并影响模型后验,提出需验证重参数化以避免错误合并模型。
AI 中文摘要
大型语言模型贝叶斯(LLB)方法通过从语言模型中采样候选概率程序、对每个程序运行近似推理,并以与指数化证据下界成比例的权重进行平均,来回答一个非正式的建模问题。我们证明,这种权重取决于模型的书写方式。对数边际似然对重参数化具有不变性,而证据下界则不具备这一性质。在八个学校数据集上,中心化与非中心化程序是相同的测度,差异为$5.7\times10^{-14}$,但它们的权重却相差$6.1\times$;重要性加权仅能将此差异降至$2.2\times$,而重现LLB实际运行的推理(即与后验矩匹配的全协方差高斯分布)后,在八个学校数据集上仍留有$1.9\times$的差异,在64维情况下则达到$8.9\times$。在不同似然族、维度和漏斗严重程度下,差异最大可达$31.9\times$,并且符号会发生反转,因此没有任何单一的书写方式在所有情况下都更优。该差异会使贝叶斯因子相对于八个自然竞争模型发生反转,并且对模型后验以及任何下游目标引入的误差,通过已知的尖锐希尔伯特距离界,由界短缺的离散度$\Delta$控制。在来自六个语言模型的360个程序中,参数化的书写方式从0%到100%的中心化程度不等,并且在单个模型内是稳定的。统计上检测等价程序在实用样本预算下可能错误地合并真正不同的模型;而验证我们自己生成的重参数化则不会出现此问题,并关闭了这一漏洞。
英文摘要
Large Language Bayes (LLB) answers an informal modelling question by sampling candidate probabilistic programs from a language model, running approximate inference on each, and averaging them with weights proportional to an exponentiated evidence bound. We show that this weighting depends on how a model is written. The log marginal likelihood is invariant to reparameterisation; the evidence bound is not. On eight schools the centered and non-centered programs are the same measure to $5.7\times10^{-14}$, yet their weights differ by $6.1\times$; importance weighting reduces this only to $2.2\times$, and reproducing the inference LLB actually runs, a full-covariance Gaussian matched to the posterior moments, still leaves $1.9\times$ on eight schools and $8.9\times$ in $64$ dimensions. Across likelihood families, dimensions and funnel severities the discrepancy reaches $31.9\times$ and reverses sign, so no single writing is uniformly preferable. It inverts Bayes factors against eight natural competitors, and the induced error in the model posterior, and in any downstream target, is controlled by the spread $Δ$ of the bound shortfalls through a known sharp Hilbert-distance bound. Across $360$ programs from six language models the parameterisation written ranges from $0\%$ to $100\%$ centered and is stable within a model. Detecting equivalent programs statistically can falsely merge genuinely different models at practical sample budgets; verifying reparameterisations we generate ourselves cannot, and closes the window.