AI 中文总结
研究分层模型中 PSIS-LOO 失败问题,提出用 Gelman-Pardoe 池化因子等预测失败折叠,采用集成重要性抽样修正,即 RB-LOO 方法,通过基纤维舒尔分解实现两级分类,提高交叉验证可靠性,改变决策结果。
AI 中文摘要
对于分层模型,帕累托平滑重要性抽样留一法交叉验证(PSIS-LOO)在随机效应坐标由数据驱动且其组较小时的折叠上会失败。我们表明,Gelman-Pardoe 池化因子和结构杠杆率可从模型结构和组大小预测这些折叠,无需形成重要性权重。在高斯线性混合模型中,杠杆率简化为组大小,给出一个设计时的映射,其在分离失败折叠(\(\hat{k}>0.7\))时的 AUC 为 0.96;在重复逻辑广义线性混合模型中,拟合后的无权重预测器的 AUC 为 0.81。修正方法是集成重要性抽样:边缘化随机效应块,仅对基础参数进行重要性抽样。这虽不是新方法,但我们贡献了其针对随机截距广义线性混合模型的观测级特化:一个解析高斯降阶和针对伯努利、二项式和泊松响应的一维求积,打包为一个即插即用的 rb_loo(fit)。与精确重拟合相比,这种边缘化估计器(RB-LOO)在单例繁重的逻辑广义线性混合模型上比矩匹配精确 3 倍。在有 97 个失败折叠的过度分散计数数据上,矩匹配留下 37 个未校正的,且不比原始 PSIS-LOO 更精确,而 RB-LOO 无需成本就能重现 82 分钟的精确重拟合(elpd RMSE 0.04)。误差会改变决策:与负二项式模型相比,PSIS-LOO 报告支持更复杂模型的决定性证据(\(z = 4.9\)),reloo 报告显著证据(\(z = 3.4\)),而 RB-LOO 重现的精确分析发现两者无差异(\(z = 1.0\))。基纤维舒尔分解将案例删除影响分为一个垂直(池化)项,它控制 PSIS-LOO 失败的位置,以及一个水平(方差分量)项,它控制 RB-LOO 自身紧张的位置,给出一个两级分类法,在仅重新拟合少数需要的折叠时就能恢复精确答案。
英文摘要
For hierarchical models, Pareto-smoothed importance-sampling leave-one-out cross-validation (PSIS-LOO) fails on the folds where a random-effect coordinate is data-driven and its group is small. We show that the Gelman-Pardoe pooling factor and structural leverage predict these folds from model structure and group sizes, without forming importance weights. In Gaussian linear mixed models the leverage reduces to group size, giving a design-time map that separates the failing ($\hat{k}>0.7$) folds with AUC 0.96; across replicated logistic GLMMs the post-fit, weight-free predictor reaches AUC 0.81. The cure is integrated importance sampling: marginalise the random-effect block and importance-sample only the base parameters. This is not new, but we contribute its observation-level specialisation for random-intercept GLMMs: an analytic Gaussian downdate and a 1-D quadrature for Bernoulli, binomial and Poisson responses, packaged as a drop-in rb_loo(fit). Against exact refits, this marginalised estimator (RB-LOO) is $3\times$ more accurate than moment matching on singleton-heavy logistic GLMMs. On overdispersed count data with 97 failing folds, moment matching leaves 37 uncorrected and is no more accurate than raw PSIS-LOO, while RB-LOO reproduces the 82-minute exact refit (elpd RMSE 0.04) at no cost. The error changes decisions: against a negative-binomial model, PSIS-LOO reports decisive evidence ($z=4.9$) and reloo reports significant evidence ($z=3.4$) for the more complex model, where an exact analysis, reproduced by RB-LOO, finds the two indistinguishable ($z=1.0$). A base-fiber Schur decomposition splits case-deletion influence into a vertical (pooling) term that governs where PSIS-LOO fails and a horizontal (variance-component) term that governs where RB-LOO is itself strained, giving a two-level triage that recovers the exact answer while refitting only the few folds that need it.
Comments16 pages, 5 figures