发表机构
Department of Statistics, Stockholm University(斯德哥尔摩大学统计系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文指出标准蒙特卡洛估计LPDS在异常值下会严重偏差,提出用完整与训练数据对数边际似然之差估计LPDS,更高效且作为PSIS-LOO的更好后备方案。
AI 中文摘要
对数预测密度分数(LPDS)及其相关估计量(用于估计期望对数预测密度(ELPD))已成为贝叶斯模型比较的事实标准。LPDS是一种自然的样本外预测密度性能度量,它考虑了参数不确定性,对先验的依赖程度低于边际似然,并且可以使用来自后验抽样的标准蒙特卡洛方法轻松地顺序计算。然而,我们的文章表明,当测试数据存在异常值时,LPDS的标准蒙特卡洛估计可能极其不稳定且严重有偏,导致误导性结论和次优模型选择。将LPDS估计为完整数据集的对数边际似然与训练数据的对数边际似然之差,被证明是一种显著更高效的方法。该估计器还被证明是对于已被Stan中广泛使用的PSIS-LOO模型比较框架标记为有问题的观测值,一种更有效的后备方案。
英文摘要
The log predictive density score (LPDS) and related estimators of the expected log predictive density (ELPD) have become the de facto standard for Bayesian model comparison. The LPDS is a natural out-of-sample measure of predictive density performance that accounts for parameter uncertainty, is less dependent on the prior than the marginal likelihood, and is easily computed sequentially using standard Monte Carlo from posterior draws. Our article demonstrates, however, that the standard Monte Carlo estimate of the LPDS can be extremely variable and severely biased when the test data has outliers, leading to misleading conclusions and suboptimal model choices. Estimating the LPDS as the difference between the log marginal likelihood of the full dataset and the training data is shown to be a dramatically more efficient approach. This estimator is also shown to be a much more effective fallback for observations that have been flagged as problematic by the widely used PSIS-LOO model comparison framework in Stan.
Comments29 pages, 13 figures