发表机构
Mofid Securities(莫菲德证券)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对协变量偏移下加权保形预测成本高的问题,提出草图校准方法,通过压缩协变量降低校准成本并控制覆盖率泄漏,实验显示其显著减少无限预测集。
AI 中文摘要
加权保形预测通过使用目标协变量与源协变量之间的似然比对校准分数进行重新加权,从而修正协变量偏移。其成本随两个协变量分布之间的卡方散度增长,通常随偏移规模呈指数增长,且其中很大一部分成本可能用于对响应无影响的偏移方向。我们提出草图校准:使用压缩协变量 $Z=T(X)$ 的比进行加权保形预测,该比用于权重,也可选择用于分数。压缩永远不会增加依赖于偏移的校准成本,且目标覆盖率至少为 $1-\alpha-\Delta_T$,其中泄漏 $\Delta_T$ 衡量被丢弃的偏移中有多少以响应或分数(给定 $Z$)的分布变化形式重新出现。泄漏是草图纤维上被丢弃偏移与响应之间的协方差;当草图是充分的或保留偏移时,泄漏消失,且极小极大构造表明,对于固定分数,任何阈值规则都无法避免泄漏。在异方差模拟中,无加权校准失败,一维草图将无限预测集的比例从 $37.8\\%$ 降至 $8.0\\%$,平均覆盖率为 $92.3\\%$(目标为 $90\\%$);非线性实验和真实数据实验展示了学习草图何时成功及何时失败。
英文摘要
Weighted conformal prediction corrects for covariate shift by reweighting calibration scores with the likelihood ratio between target and source covariates. Its cost grows with the chi-square divergence between the two covariate laws, typically exponentially in the size of the shift, and much of it can be paid for shift in directions that do not affect the response. We propose sketched calibration: weighted conformal prediction with the ratio of compressed covariates $Z=T(X)$, used in the weights and optionally in the score. Compression never increases the shift-dependent calibration cost, and target coverage is at least $1-α-Δ_T$, where the leakage $Δ_T$ measures how much of the discarded shift reappears as a change in the law of the response, or of the score, given $Z$. The leakage is a covariance between the discarded shift and the response on the fibres of the sketch; it vanishes when the sketch is sufficient or retains the shift, and a minimax construction shows that no threshold rule for a fixed score can avoid it. In a heteroscedastic simulation where unweighted calibration fails, a one-dimensional sketch cuts the fraction of infinite prediction sets from $37.8\%$ to $8.0\%$ at mean coverage $92.3\%$ for a $90\%$ target; nonlinear and real-data experiments show when learned sketches succeed and when they fail.