arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可复现的共形预测

Replicable Conformal Prediction

Marios Papamichalis, Regina Ruane, Theofanis Papamichalis

arXiv 2608.23638首次发表:更新:

发表机构

Yale University; University of Pennsylvania; The Wharton School(耶鲁大学; 宾夕法尼亚大学; 沃顿商学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对共形预测的不稳定性问题,提出通过共享随机种子并将校准阈值舍入到共享网格的方法,实现分类器可复现且保留覆盖保证,经多类实验验证符合理论且能阻止博弈。

AI 中文摘要

两名分析人员在独立样本上对同一预测模型进行校准时,每次都会部署不同的预测集,因为校准阈值继承了数据的随机性。当部署需要跨站点审计、缓存或批准时,这种不稳定性成本很高:没人能验证两次校准是否产生了相同的对象。我们提出两个问题:独立校准何时能产生相同的分类器,以及这种一致性的代价是什么?完美的一致性是不可能的,因为几乎总是返回一个固定答案的过程无法对所有分布保持有效性,而通过共享随机性实现的精确一致性会迫使该过程忽略其数据。共享单个随机种子并将校准阈值向上舍入到一个粗糙的共享网格可解决这一矛盾:部署的分类器在分析人员之间以任意期望的概率变得相同,覆盖保证得以保留,代价是集合大小和校准数据的量化增加。匹配的下界表明,没有阈值方法能付出更少的代价,且该方法的一个调优常数会渐近消失。没有任何共享种子时,固定网格仍会将所有分析人员限制在两个相邻的分类器中,且没有方法能做得更好。可复现性还能阻止博弈:从多次重新校准中选择最有利的结果几乎不会改变可复现分类器,而相同的选择会使标准共形预测出现隐性覆盖不足。在真实ImageNet输出、四医院站点划分以及四个语言模型族上的实验与理论相符,包括测量的样本成本边界。

英文摘要

Two analysts who calibrate the same predictive model on independent samples will deploy different prediction sets every time, because the calibration threshold inherits the randomness of the data. Wherever deployments must be audited, cached, or approved across sites, this instability is costly: no one can verify that two calibrations produced the same object. We ask two questions: when can independent calibrations yield the identical classifier, and what must that agreement cost? Perfect agreement is impossible, since a procedure that almost always returns one fixed answer cannot remain valid for every distribution, and exact agreement through shared randomness forces the procedure to ignore its data. Sharing a single random seed and rounding the calibrated threshold up to a coarse shared grid resolves the tension: the deployed classifier becomes identical across analysts with any desired probability, coverage guarantees survive, and the price is a quantified increase in set size and calibration data. Matching lower bounds show that no threshold method can pay less, and the method's one tuning constant vanishes asymptotically. Without any shared seed, a fixed grid still confines all analysts to two adjacent classifiers, and no method does better. Replicability also blocks gaming: selecting the most favorable of many recalibrations barely moves a replicable classifier, while the same selection silently undercovers standard conformal prediction. Experiments on real ImageNet outputs, a four-hospital site split, and four language-model families match the theory, including the measured sample-cost frontier.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑