发表机构
Chung-Ang University; Yonsei University(中央大学; 延世大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对测试时提示调优(TPT)的熵最小化目标导致模型校准退化的问题,提出新的对齐目标并结合置信度温度缩放,在多基准上实现了最优准确率与校准性能的提升。
AI 中文摘要
测试时提示调优(Test-time Prompt Tuning,TPT)已成为一种强大的范式,通过在多个增强视图上的熵最小化(Entropy Minimization,EM)为每个测试样本优化提示。然而,我们发现基于标准EM的自适应存在一个局限:它本质上会驱使模型做出过度自信的预测,而忽略样本特定的不确定性,导致校准性能显著下降。为解决这些局限,我们提出了一个新目标,该目标用交叉熵将原始视图的预测与从增强视图得到的目标分布对齐,同时对抗性地纳入目标分布的熵以捕获样本特定的不确定性。此外,为更好地构建该目标分布,我们根据每个增强视图预测的置信度对其应用感知置信度的温度缩放,对置信的预测进行锐化,对不确定的预测进行软化。该公式使模型仅在目标分布可靠时才提高置信度,而当目标分布反映模糊或冲突的增强视图预测时则保留不确定性。在不同基准上的大量实验表明,我们的方法不仅达到了最先进的准确率,还显著提升了模型校准性能。
英文摘要
Test-time prompt tuning (TPT) has emerged as a powerful paradigm, refining prompts for each test sample via entropy minimization (EM) over multiple augmented views. However, we identify a limitation in the standard EM-based adaptation: it inherently drives the model toward overconfident predictions disregarding sample-specific uncertainty, leading to significant calibration degradation. To address these limitations, we propose a new objective that replaces the conventional EM loss by aligning the original-view prediction with a target distribution derived from augmented views via cross-entropy, while adversarially incorporating the entropy of the target distribution to capture sample-specific uncertainty. Furthermore, to better construct this target distribution, we apply confidence-aware temperature scaling to each augmented-view prediction according to its confidence, sharpening confident predictions while softening uncertain ones. This formulation allows the model to increase confidence only when the target distribution is reliable, while preserving uncertainty when it reflects ambiguous or conflicting augmented-view predictions. Extensive experiments across diverse benchmarks demonstrate that our approach not only achieves state-of-the-art accuracy but also significantly improves model calibration.
Comments9 pages