发表机构
University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究分布偏移下类条件覆盖的标签复杂性,指出无标签方法难以兼顾有效性和效率,精确了恢复每类有效性的成本,通过骨架动作识别等案例研究表明相关现象在多模态基准上存在。
AI 中文摘要
许多识别系统的标准评估中存在分布偏移,因为基准在训练和测试分割中设置了不相交的条件。在这种偏移下,分割共形预测使边际覆盖率接近标称水平,而每类覆盖率却悄然失效。我们刻画了恢复每类有效性的成本。首先,存在一个不可能的情况:一旦偏移同时作用于协变量和标签,目标类条件得分律就无法从源标签和未标记的目标样本中识别出来,所以没有无标签方法能同时实现有效且高效的每类覆盖率。其次,我们精确了成本:仅每类有效性只需每类少量目标标签,而要同时实现有效性和每类效率所需的标签数量随着效率容差的平方反比和类数的对数增长,且有匹配的上下界。第三,在评估的预测驱动推理家族中,即使在无界未标记目标池上最有利地使用分类器自己的伪标签,在覆盖率崩溃时效率最多提高一个小常数因子。骨架动作识别是我们的真实数据案例研究。仅使用源标签进行每类校准可在偏移保持边际覆盖率时恢复大部分每类差距,且在边际覆盖率本身崩溃时停止起作用。三种严重程度不断增加的真实偏移追踪了这个边界,并且在自然图像损坏基准上也出现了同样的崩溃和恢复,超越了任何单一模态。
英文摘要
Prediction sets can make deployed classifiers safer by returning several plausible labels when a single prediction is uncertain. Their value depends on classwise reliability: average coverage can meet its target while rare or difficult classes fail repeatedly. This concern is sharper after distribution shift, when calibration labels come from a source environment but reliability is needed on the target. We ask what labeled source data and unlabeled target inputs reveal about class-conditional prediction sets, and when target labels are necessary. Under unrestricted joint shift, two target laws can produce the same observable data while requiring different classwise thresholds; any label-free rule covering both must enlarge its sets on one law. We give a labeled target audit that estimates the missing quantiles with a simultaneous guarantee. Probability-scale error is invariant to increasing score transformations, and threshold recovery follows under local regularity. At fixed confidence, achieving threshold tolerance $\varepsilon$ with fixed, nonadaptive class-stratified labeled pairs has total complexity $Θ(K\varepsilon^{-2}\log K)$, or $Θ(\varepsilon^{-2}\log K)$ labels per class under equal allocation. Class imbalance creates a separate acquisition cost; for foreground class probabilities of order $1/K$, the mixed-stream label complexity is also $Θ(K\varepsilon^{-2}\log K)$ at fixed confidence. Experiments on action-recognition and image shifts show that marginal coverage can conceal severe class failures and that source classwise calibration depends on the shift. The results connect the information available at deployment to the target labels needed for useful class-conditional prediction.
Comments35 pages, 3 figures; includes appendices