arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25138stat.MEcs.CVcs.LG

校准计数重用:有效性不决定效率

Calibration Count Reuse: Validity Does Not Determine Efficiency

Rudra Chopra

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨校准计数重用中有效性与效率的独立性,提出有效性准则并证明效率无排序,通过扩展规则研究和图像实验揭示覆盖不足,强调基准结果非普遍保证。

中文摘要 AI 辅助

校准计数重用提出了独立的有效性和效率问题。我们针对一般的计数依赖非一致性评分给出了一个有效性准则:将计数从另一个类别转移到被评分类别不得提高其一致性。一个留一自排除的完全共形参考证明了该准则,无需归一化或保持同类分数顺序。对于常见的可分离变换,当$K\alpha \geq 1$时,通用可交换有效性与计数非递减性等价;归一化乘法权重满足互补的非递增条件。加性惩罚在所述信息限制下被覆盖。效率没有平行的排序:两种独立同分布构造使同一有效规则在覆盖率不变的情况下,期望大小分别改善或恶化。一项扩展的55规则研究发现,选定的实时计数规则相对于均匀权重没有解决的优势。图像研究识别出在独立同分布重采样下的覆盖不足,包括在数值收敛时。另外,在发布的Conf-OT流水线在其DTD和Aircraft基准子集上的执行,在固定分层计数下产生接近标称的中位覆盖率。原生结果与独立同分布分析分开报告,不将基准观察视为普遍保证。研究结果区分了有效性、分类器置信度、数值收敛和群体特定效率。

英文摘要

Calibration count reuse raises separate validity and efficiency questions. We give a validity criterion for general count-dependent nonconformity scores: transferring one count from another class to the scored class must not improve its conformity. A leave-self-out full conformal reference proves the criterion without requiring normalization or preservation of same-class score order. For a common separable transformation, universal exchangeable validity is equivalent to being nondecreasing in the count, provided $Kα\geq 1$; normalized multiplicative weights obey the complementary nonincreasing condition. Additive penalties are covered under the stated information restrictions. Efficiency has no parallel ordering: two iid constructions make the same valid rule improve or worsen expected size at unchanged coverage. An expanded 55-rule study finds no resolved advantage from selected live-count rules over uniform weights. Image studies identify undercoverage under iid resampling, including at numerical convergence. Separately, execution of the released Conf-OT pipeline on its DTD and Aircraft benchmark subsets produces near-nominal median coverage under fixed stratified counts. The native results are reported separately from the iid analyses, without treating a benchmark observation as a universal guarantee. The findings separate validity, classifier confidence, numerical convergence, and population-specific efficiency.

补充信息

↑