AI 中文总结
本研究提出面向预测(PrO)推理的有限维对偶形式与近似PrO后验的有限样本超额预测风险界,通过示例验证其在模型误设定场景下的有效性。
AI 中文摘要
面向预测(PrO)推理通过选择模型参数上的分布,优化应用于诱导预测分布的评分规则,同时结合与参考分布的散度惩罚,以此量化不确定性。由于在对模型密度取平均后应用评分规则,PrO推理以预测性能为目标,同时考虑了模型误设定问题。本研究聚焦于带有一般φ-散度正则化的对数评分,贡献分为两部分:其一,推导了PrO推理的有限维对偶形式,对于n个观测值,对偶问题有n+1个变量;建立了零对偶间隙准则及关联原始与对偶解的最优性条件,当原始与对偶解存在时,这些条件可得到PrO后验的半解析表示,以及用于评估数值解准确性的验证指标;对于Kullback-Leibler正则化,后验具有指数形式。其二,推导了近似PrO后验的有限样本超额预测风险界,该界将抽样波动、散度预算下的近似误差、正则化误差及数值优化误差分离开来;该结果甚至适用于参考分布的有限散度模型参数上的概率分布无法达到基准预测风险的情况。本研究采用可精确求解的分类示例表明,预测风险收敛可意味着收敛到唯一的预测分布,即便参数分布在原始参数空间上没有弱极限;该示例还显示,不同的φ-散度需要不同的正则化方案。最后,本研究以误设定的高斯位置混合示例说明对偶计算、原始解恢复及数值准确性检查。
英文摘要
Predictively oriented (PrO) inference quantifies uncertainty by selecting a distribution over model parameters to optimize a scoring rule applied to the induced predictive distribution, together with a divergence penalty from a reference distribution. By applying the scoring rule after averaging model densities, PrO inference targets predictive performance, accounting for model misspecification. We focus on the logarithmic score with general $ϕ$-divergence regularization. Our contributions are twofold. First, we derive a finite-dimensional dual formulation of PrO inference. For $n$ observations, the dual problem has $n+1$ variables. We establish zero-duality-gap criteria and optimality conditions that relate the primal and dual solutions. When primal and dual solutions exist, these conditions yield a semi-analytical representation of the PrO posterior and certificates for assessing the accuracy of numerical solutions. For Kullback--Leibler regularization, the posterior has an exponential form. Second, we derive a finite-sample excess predictive-risk bound for approximate PrO posteriors that separates sampling fluctuation, approximation under a divergence budget, regularization, and numerical optimization error. The result applies even when the benchmark predictive risk is not attained by any probability distribution over the model parameters having finite divergence from the reference distribution. We use an exactly solvable categorical example to show that predictive-risk convergence can imply convergence to a unique predictive distribution even though the parameter distributions have no weak limit on the original parameter space. The example also shows that different $ϕ$-divergences can require different regularization schedules. We conclude with a misspecified Gaussian location-mixture example that illustrates the dual computation, primal recovery, and numerical accuracy checks.