arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

广义线性混合效应模型的先验摊销上下文贝叶斯推断

Prior-Amortized In-Context Bayesian Inference for Generalized Linear Mixed-Effects Models

Alex Kipnis, Marcel Binz, Eric Schulz

arXiv 2609.24422首次发表:更新:

发表机构

Institute for Human-Centered AI, Helmholtz Munich(慕尼黑亥姆霍兹中心人本人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出metabeta,一种预训练神经网络,通过先验摊销和上下文学习实现GLMMs的贝叶斯推断,无需MCMC,速度比NUTS快两到三个数量级,且保持推断准确性。

AI 中文摘要

分层数据在实证科学中无处不在,最常用的分析方法是广义线性混合效应模型(GLMMs)。GLMMs的贝叶斯推断能提供校准的不确定性,但需要MCMC;No-U-Turn采样器(NUTS)是黄金标准,但速度慢,且必须对每个新数据集、模型和先验从头开始。我们提出了metabeta,一个用于GLMMs先验摊销上下文贝叶斯推断的预训练神经网络。与以往在训练时固定先验的神经后验估计器不同,metabeta在测试时接受先验族和超参数作为输入,实现零样本泛化。两个集合变换器和条件归一化流反映了后验的两级结构(跨组共享的全局参数,每组的局部参数)。该模型在数百万个涵盖连续、二值和计数结果的逼真模拟数据集上训练。默认情况下,流后验通过独立Metropolis-Hastings针对未归一化后验进行细化,因此其正确性依赖于采样器而非网络;这实现了比NUTS快两到三个数量级的免调参推断。或者,流可以热启动NUTS,在显著提高速度和稳定性的同时,给出几乎相同的推断。在具有真实参数的控制基准上,metabeta在参数恢复、校准和样本外预测方面与NUTS匹配。在分布外的真实数据集上,其后验在所有参数类型上与NUTS的后验紧密匹配,并且在错误指定的似然和先验、分布外预测变量、共线设计和数据稀缺情况下保持忠实。该模型开源且开放权重,因此可立即部署。

英文摘要

Hierarchical data is ubiquitous in the empirical sciences and is most commonly analyzed with generalized linear mixed-effects models (GLMMs). Bayesian inference for GLMMs yields calibrated uncertainty but requires MCMC; the No-U-Turn Sampler (NUTS) is the gold standard but is slow and must restart from scratch for every new dataset, model and prior. We introduce metabeta, a pretrained neural network for prior-amortized in-context Bayesian inference over GLMMs. Unlike previous neural posterior estimators that fix the prior at training time, metabeta accepts prior families and hyperparameters as inputs at test time, enabling zero-shot generalization. Two set transformers and conditional normalizing flows mirror the posterior's two-level structure (global parameters shared across groups, local parameters per group). The model is trained on millions of realistic simulated datasets spanning continuous, binary, and count outcomes. By default, the flow posterior is refined by Independence Metropolis-Hastings against the unnormalized posterior, so its correctness rests on the sampler rather than the network; this yields tuning-free inference two to three orders of magnitude faster than NUTS. Alternatively, the flow can warm-start NUTS, giving nearly identical inference with substantially increased speed and stability. On controlled benchmarks with ground-truth parameters, metabeta matches NUTS in parameter recovery, calibration and out-of-sample prediction. On out-of-distribution real datasets, its posteriors closely match those of NUTS across all parameter types, and they remain faithful under misspecified likelihoods and priors, out-of-distribution predictors, collinear designs, and data-poor regimes. The model is open-source and open-weights and thus immediately deployable.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑