arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05945cs.CV

重视零样本不确定性:面向测试时自适应视觉-语言模型的保守校准

Respect Your Zero-Shot Uncertainty: Conservative Calibration for Test-Time-Adapted Vision-Language Models

Jingyan Jiang, Yaru Sun, Xiao Chen, Jiazhen Huang, Caiting Li, Zhijian He, Yin Chen, Pingting Hao

首次发表
浏览论文内容

中文总结 AI 辅助

针对测试时自适应(TTA)视觉-语言模型校准度下降的问题,提出零样本锚定熵校准(ZAEC)方法,通过最小温度缩放恢复锐化预测的零样本熵,在多模型多数据集上实现了更低的期望校准误差(ECE)。

中文摘要 AI 辅助

测试时自适应(TTA)可提升视觉-语言模型在分布偏移下的识别准确率,但常会降低校准度,使预测置信度对下游决策不可靠。许多现有的无标签校准方法要么与提示优化耦合,要么依赖仅能粗略刻画预测分布的 logit 范围统计量。我们发现,即使 Top-1 预测及其正确性保持不变,TTA 也可能提升置信度并降低熵,这种失效模式我们称为“预测保留型锐化”。在多种 TTA 方法和基准上,与配对的零样本预测相比,更大的熵降低与期望校准误差(ECE)的更大提升相关;在熵降低的样本上,置信度提升也往往超过准确率提升。基于这些发现,我们提出零样本锚定熵校准(ZAEC),这是一种无标签的事后方法,它将零样本熵作为样本特定的不确定性参考,通过最小温度缩放选择性恢复锐化预测的零样本熵,同时保持所有其他预测不变。该方法无需带标签的校准数据或学习参数,且保留类别排名和分类准确率。在五种 TTA 方法和 15 个数据集上,ZAEC 在 ViT-B/16 上实现了最低的事后宏平均 ECE,在 RN50 上也取得了一致的提升。

英文摘要

Test-time adaptation (TTA) can improve the recognition accuracy of vision-language models under distribution shift, but often degrades calibration, making predictive confidence unreliable for downstream decision-making. Many existing label-free calibration approaches are either coupled to prompt optimization or rely on logit-range statistics that provide only a coarse characterization of the predictive distribution. We show that TTA can increase confidence and reduce entropy even when the top-1 prediction and its correctness remain unchanged, a failure mode we term prediction-preserving sharpening. Across diverse TTA methods and benchmarks, larger entropy reductions relative to paired zero-shot predictions are associated with greater increases in Expected Calibration Error (ECE). On entropy-reduced samples, confidence gains also tend to exceed accuracy gains. Based on these findings, we propose Zero-Shot-Anchored Entropy Calibration (ZAEC), a label-free post-hoc method that uses zero-shot entropy as a sample-specific uncertainty reference. ZAEC selectively restores the zero-shot entropy of sharpened predictions through minimal temperature scaling while leaving all other predictions unchanged. It requires no labeled calibration data or learned parameters and preserves class rankings and classification accuracy. Across five TTA methods and 15 datasets, ZAEC achieves the lowest post-hoc macro-average ECE on ViT-B/16, with consistent gains on RN50.

↑