arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39934cs.LGcs.CV

面向域泛化的可靠性感知检查点选择

Reliability-Aware Checkpoint Selection for Domain Generalization

  • Shenzhen University(深圳大学)
  • Xiamen University(厦门大学)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • Macao Polytechnic University(澳门理工大学)
  • Fudan University(复旦大学)
  • Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)

机构由 AI 辅助整理,请以论文原文为准。

Jinshi Liu, Jiahao Li, Pan Liu, Yanfeng Li, Rui Qian, Zhao Tong, Yue Sun, Tao Tan

AI总结:

本文提出一种可靠性感知的检查点重选方法,在源准确率约束下利用可靠性指标排序,无需目标数据即可提升域泛化中目标域的概率质量。

AI中文摘要:

在域泛化中,检查点选择通常依赖于源验证准确率,然而所选检查点未必能在未见的目标域上提供可靠的预测概率。源域与目标域之间的分布偏移可能改变准确率排名,而准确率本身并不能衡量预测概率的质量。我们在固定的训练轨迹中发现了一个经验性的选择机会:在具有接近最优源准确率的检查点中重新选择,可以在平均目标准确率发生较小观测变化的同时,提升平均目标概率质量。我们研究了准确率约束下的可靠性选择(AC),该方法保留在最佳源验证准确率容差范围内的检查点,并按源可靠性对其进行排序。我们的参考规则使用$D_\infty$聚合集合内归一化负对数似然(NLL)和逐类校准误差(CwECE)。AC不使用任何目标数据,且无需额外训练或权重平均。我们在三个基准上评估了五种域泛化训练算法,使用PACS来制定目标,并采用0.5个百分点的容差。在对360个OfficeHome和TerraIncognita运行的探索性聚合比较中,参考规则相对于Source-Acc将平均目标软箱平方间隙ECE和CwECE分别降低了0.240%和0.182%,并将NLL降低了0.030。平均目标准确率变化了+0.213个百分点。这些结果揭示了可靠性感知重新选择的机会,而联合目标相对于单目标排序的额外收益仍未得到解决。

英文摘要:

Checkpoint selection in domain generalization often relies on source-validation accuracy, yet the selected checkpoint need not provide reliable probabilities on unseen target domains. Source-target distribution shifts can alter accuracy rankings, while accuracy alone does not measure predictive probability quality. We identify an empirical selection opportunity within fixed training trajectories: reselecting among checkpoints with near-optimal source accuracy can improve mean target probability quality with small observed changes in mean target accuracy. We study accuracy-constrained reliability selection (AC), which retains checkpoints within a tolerance of the best source-validation accuracy and ranks them by source reliability. Our reference rule aggregates within-set normalized negative log-likelihood (NLL) and class-wise calibration error (CwECE) using $D_\infty$. AC uses no target data and requires neither additional training nor weight averaging. We evaluate five domain generalization training algorithms on three benchmarks, using PACS to develop the objectives and a 0.5-percentage-point tolerance. In exploratory aggregation comparisons on 360 OfficeHome and TerraIncognita runs, the reference rule reduces mean target soft-bin squared-gap ECE and CwECE by 0.240% and 0.182%, respectively, and NLL by 0.030 relative to Source-Acc. Mean target accuracy changes by +0.213 percentage points. These results identify opportunities for reliability-aware reselection, while the additional benefit of joint over single-objective ranking remains unresolved.

补充信息

↑