arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11349cs.LGstat.ME

样本高效的生成式共形预测

Sample-Efficient Generative Conformal Prediction

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Minxing Zheng, Shixiang Zhu

AI总结:

针对生成式共形预测采样效率低的问题,提出CASA方法,通过自适应分配样本在相同预算下得到更小的不确定性集合,还可改善条件覆盖率并与现有方法互补。

AI中文摘要:

生成式共形预测从条件生成器的样本构建不确定性集合,仅当样本能很好地表示响应分布时才高效。这需要大量样本,而每个样本可能成本高昂,如大型扩散模型和科学模拟器中的情况,因此必须高效利用采样预算。现有方法在每个输入处抽取相同数量的样本,在响应分布简单的地方浪费样本,在复杂的地方采样不足,这会扩大集合并导致这些输入覆盖不足。我们提出CASA(Conformal Adaptive Sample Allocation,共形自适应样本分配),其表征额外样本的边际价值,并在输入间分配样本,以在边际覆盖率和平均采样预算的约束下最小化期望集合大小。理论分析表明,在相同预算下,自适应分配比固定样本数产生的集合更小:错过一个模式会迫使半径跨越模式间的间隙,即使是最优半径也无法弥补这一点。在合成任务和真实任务上,CASA在相同预算下产生显著更小的集合,常能改善条件覆盖率,且可与现有的半径自适应方法互补。

英文摘要:

Generative conformal prediction builds uncertainty sets from samples of a conditional generator, which are efficient only when the samples represent the response distribution well. This can require many samples, each of which can be costly, as in large diffusion models and scientific simulators, so the sampling budget must be used efficiently. Existing methods draw the same number of samples at every input, wasting samples where the response distribution is simple and undersampling where it is complex, which inflates sets and leaves those inputs under-covered. We propose CASA (Conformal Adaptive Sample Allocation), which characterizes the marginal value of an additional sample and allocates samples across inputs to minimize the expected set size subject to marginal coverage and an average sampling budget. Theoretical analysis shows that adaptive allocation yields smaller sets than a fixed count at the same budget: a missed mode forces a radius that spans the gap between modes, and even oracle radius cannot compensate for it. On synthetic and real tasks, CASA produces substantially smaller sets at the same budget, often improves conditional coverage, and complements existing radius-adaptive methods.

↑