arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

测试时自适应何时起作用、造成损害或无效:针对CIFAR-10-C的条件级研究

When Test-Time Adaptation Helps, Harms, or Becomes Inactive: A Condition-Level Study on CIFAR-10-C

Sreeja Guha Majumdar, Aratrika Saha

arXiv 2608.22233首次发表:更新:

发表机构

Heritage Institute of Technology(遗产技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对CIFAR-10-C开展条件级分析,对比三种测试时自适应策略与源模型的性能,发现自适应存在失效场景,揭示可靠性过滤会影响EATA的表现,强调需开展条件级评估。

AI 中文摘要

测试时自适应(TTA)旨在通过使用未标记的测试数据对源模型进行自适应,提升模型在分布偏移下的鲁棒性。尽管TENT、EATA等方法在损坏数据上已展现出性能提升,但整体准确率可能掩盖自适应失效或收益甚微的条件。我们针对完整的CIFAR-10-C基准(涵盖15种损坏类型和5个严重程度等级),将三种TTA策略——批量归一化统计量自适应(BN-Adapt)、熵最小化自适应(TENT)、可靠性过滤自适应(EATA的限定范围重实现)——与未自适应的源模型进行受控对比。三种方法的平均准确率较源模型提升了12.2至13.3个百分点(Wilcoxon符号秩检验p值<10^-12),但每种方法在8.0%至9.3%的条件下性能均逊于源模型,失效集中在源模型已接近性能上限的低严重程度损坏中,尤其是亮度、雾、对比度和散焦模糊场景。我们进一步发现,EATA与无梯度的BN-Adapt基线的平均绝对差为0.09个百分点,而与TENT的平均绝对差为1.08个百分点,这表明可靠性过滤可大幅限制有效自适应,使EATA的表现更接近批量归一化统计量基线,而非熵最小化方法。这些结果表明,仅靠整体准确率会掩盖系统性的TTA失效模式,推动开展条件级评估以明确自适应何时起作用、造成损害或变得无效。

英文摘要

Test-time adaptation (TTA) aims to improve model robustness under distribution shift by adapting a source model using unlabeled test data. Although methods such as TENT and EATA have demonstrated gains on corrupted data, aggregate accuracy can obscure the conditions under which adaptation fails or provides little benefit. We present a controlled comparison of three TTA strategies---BatchNorm-statistics adaptation (BN-Adapt), entropy-minimization adaptation (TENT), and reliability-filtered adaptation (a scoped re-implementation of EATA)---against an unadapted source model on the full CIFAR-10-C benchmark, covering 15 corruption types and 5 severity levels. All three methods improve mean accuracy over the source model by 12.2--13.3 percentage points (Wilcoxon signed-rank $p < 10^{-12}$). However, each method underperforms the source model on 8.0--9.3\% of conditions, with failures concentrated in low-severity corruptions where the source model already performs near ceiling, particularly brightness, fog, contrast, and defocus blur. We further find that EATA closely tracks the gradient-free BN-Adapt baseline, with a mean absolute difference of 0.09 percentage points, compared with 1.08 percentage points relative to TENT. This suggests that reliability filtering can substantially restrict effective adaptation, causing EATA to behave more like a BatchNorm-statistics baseline than an entropy-minimization method. These results show that aggregate accuracy alone can mask systematic TTA failure modes and motivate condition-level evaluation of when adaptation helps, harms, or becomes effectively inactive.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑