arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00928cs.LG

子类型鲁棒性不止是准确率:未见子类型分布偏移下的校准

Subtype Robustness Is Not Just Accuracy: Calibration Under Unseen Subtype Shift

Hanyu Su, Carlota Julbe i Juanola, Yibo Hu

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过多数据集与多架构的系统性研究,发现子类型偏移下模型校准失效,仅靠准确率评估子类型鲁棒性不足,需结合校准指标。

中文摘要 AI 辅助

子类型鲁棒性研究的是,当测试样本来自训练时未出现但仍属于已知粗类别的细粒度子类型时,模型是否仍能做出正确的粗粒度预测。此前研究几乎完全通过准确率来考察这一特性,本文则探究模型是否还能保持校准状态。我们针对ImageNet、BREEDS、iNaturalist和CIFAR-100四个数据集,结合五种架构开展了该问题的首次系统性研究。结果显示,在未见子类型上校准会失效:准确率下降的同时,置信度几乎未同步变化,模型在准确率降低的区域系统性过度自信;在相同的准确率损失下,通用图像损坏导致的置信度下降幅度远大于子类型偏移,说明该效应并非准确率下降的普遍结果,模型对可见退化有反应但对分类内新奇性无反应;针对可见子类型调优的重新校准可缩小但无法消除该差距,分布外分数对受影响输入的标记效果很弱。因此,子类型鲁棒性应通过校准而非仅准确率来评估。

英文摘要

Subtype robustness asks whether a model keeps the correct coarse prediction when test examples come from fine-grained subtypes absent from training but still inside a known coarse category. Prior work studies this almost entirely through accuracy. We ask whether the model also stays calibrated. We present the first systematic study of the question across ImageNet, BREEDS, iNaturalist and CIFAR-100 with five architectures. Calibration breaks down on unseen subtypes, where accuracy drops while confidence barely follows, leaving the model systematically overconfident exactly where it has become less accurate. At matched accuracy loss, generic image corruption causes a much larger drop in confidence, so the effect is not a general consequence of losing accuracy. The model reacts to visible degradation but not to in-taxonomy novelty. Recalibration tuned on seen subtypes narrows the gap but does not close it, and out-of-distribution scores flag the affected inputs only weakly. Subtype robustness should therefore be evaluated through calibration, not accuracy alone.

补充信息

↑