超越响应预测:针对潜在分布参数的共形推理
Beyond Prediction: Conformal Inference for Latent Distributional Parameters
浏览论文内容
中文总结 AI 辅助
本文提出无先验共形框架LatentCP,用于构造潜在分布参数的不确定性集合,经合成与真实野火数据集验证,其在多种场景下可保持良好覆盖率,独立调优与多级聚合可提升效率。
中文摘要 AI 辅助
许多预测问题旨在推断控制可观测响应分布的未观测实例特定参数,尽管该潜在参数无法用于历史实例和未来实例。我们开发了LatentCP,这是一种无先验的共形框架,仅使用观测到的上下文-响应对和指定的正向模型,为潜在分布参数构造不确定性集合。该方法首先在可观测响应空间中构造共形预测集合,然后根据其诱导的响应分布分配给该集合的概率保留候选潜在参数。这种逆方法提供了有限样本边际覆盖率,无需潜在校准标签、唯一逆映射或潜在混合分布的知识。由于潜在集合效率非单调依赖于响应空间的未覆盖率,我们进一步引入了多级程序,该程序聚合多个响应集的归一化不兼容得分,并使用独立调优样本选择聚合分布。在合成实验中,LatentCP在弱正向识别、观测非可识别性以及潜在异质性和多模态下保持名义潜在覆盖率,而经验贝叶斯、基于似然的和代理标签共形方法可能会严重覆盖不足。在加利福尼亚野火真实数据集上,它为潜在火灾强度产生空间自适应不确定性集合。独立调优提高了效率,而当不同响应水平包含互补信息时,多级聚合提供了额外收益,且不牺牲有效性。
英文摘要
Many prediction problems seek to infer an unobserved, instance-specific parameter that governs the distribution of an observable response, even though the latent parameter is unavailable for both historical and future instances. We develop LatentCP, a prior-free conformal framework that constructs uncertainty sets for latent distributional parameters using only observed context--response pairs and a specified forward model. The method first constructs a conformal prediction set in the observable response space and then retains candidate latent parameters according to the probability their induced response distributions assign to that set. This inversion provides finite-sample marginal coverage without requiring latent calibration labels, a unique inverse mapping, or knowledge of the latent mixing distribution. Because latent-set efficiency depends nonmonotonically on the response-space miscoverage rate, we further introduce a multilevel procedure that aggregates normalized incompatibility scores across several response sets and selects the aggregation distribution using an independent tuning sample. Across synthetic experiments, LatentCP maintains nominal latent coverage under weak forward identification, observational nonidentifiability, and latent heterogeneity and multimodality, where empirical-Bayes, likelihood-based, and proxy-label conformal methods can substantially under-cover. On a California wildfire real dataset, it produces spatially adaptive uncertainty sets for latent fire intensity. Independent tuning improves efficiency, while multilevel aggregation provides additional gains when different response levels contain complementary information, without sacrificing validity.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。