arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

模型选择需要多少标签?选择性预测的证书与预算

How Many Labels Does Model Choice Need? Certificates and Budgets for Selective Prediction

Tetsuji Kuboyama

arXiv 2609.18622首次发表:更新:

发表机构

Gakushuin University(学习院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究量化了选择性预测中模型选择所需的标签数量,提出证书大小与预算下界,并证明AUGRC选择比准确率选择需要更多标签,平均证书需56-57%的标签。

AI 中文摘要

分类器可能做出相同的预测,却需要标签来比较其选择性性能:置信度排名对相同的错误赋予不同的权重。我们针对广义风险覆盖曲线下面积(AUGRC)量化了这一需求。标签前的下界排除了不足的预算。在已知所有标签的情况下,一个覆盖线性规划限定了足以确定获胜者的最小标签数量(证书大小),对于K个候选者,该数量在K-1个标签之内。对于固定的K,独立均匀顺序和相同预测下,标签前界接近池的四分之一。当iid伯努利错误独立于顺序时,每个精确采集策略渐近地读取几乎所有标签,尽管双候选证书仅需一半。在九个数据集上的108个特征面板比较中,分歧标签解决了所有准确率选择,但未解决任何AUGRC选择。20%的预算在96种条件下被排除;证书平均需要56-57%。在十个使用预训练图像分类器的条件下,置信度分数选择读取了10,000个标签中的68-91%用于精确选择,在AUGRC容差5×10⁻⁴下为50-67%。一个精确的停止测试适用于任何采集顺序。这些结果共同将置信度排名与标签预算和认证模型比较联系起来。

英文摘要

Classifiers can make identical predictions yet require labels to compare their selective performance: confidence ranks weight the same errors differently. We quantify this requirement for the area under the generalized risk-coverage curve (AUGRC). A prelabel lower bound rules out insufficient budgets. With all labels known, a covering linear program bounds the minimum number of labels sufficient to fix the winner (the certificate size) within $K-1$ labels for $K$ candidates. For fixed $K$, independent uniform orders and identical predictions, the prelabel bound approaches one quarter of the pool. With iid Bernoulli errors independent of the orders, every exact acquisition policy reads almost all labels asymptotically, although a two-candidate certificate needs only half. Across 108 feature-panel comparisons on nine datasets, disagreement labels settle every accuracy choice but no AUGRC choice. A 20% budget is ruled out in 96 conditions; certificates need 56-57% on average. On ten conditions with pretrained image classifiers, confidence-score choice reads 68-91% of 10,000 labels for exact selection and 50-67% with AUGRC tolerance $5\times10^{-4}$. An exact stopping test works with any acquisition order. Together, these results link confidence ranks to label budgets and certified model comparison.

Comments22 pages, 13 figures, 3 tables. Includes proofs and experimental details in the main text

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑