信号路由温度缩放:低容量风险条件校准用于小验证预算
Signal-Routed Temperature Scaling: Low-Capacity Risk-Conditioned Calibration for Small Validation Budgets
浏览论文内容
中文总结 AI 辅助
针对小验证预算下事后校准器容量选择问题,提出10参数信号路由温度缩放(SRTS-BCE),在CIFAR-100等基准上以低容量匹配高容量方法性能,并揭示容量偏好随数据量变化。
中文摘要 AI 辅助
当分类器仅从几千个留出样本进行重新校准时,校准映射的容量成为一个统计设计选择,而非纯粹的结构性选择:标量映射可能欠拟合结构化的残差误校准,而高度自适应的映射则难以从如此小的数据划分中可靠估计。我们将校准目标与自适应容量分离,提出信号路由温度缩放(SRTS-BCE),这是一种10参数、保持argmax的校准器,它基于六个logit统计量对正确性风险评分进行交叉拟合,并为每个K=3风险组拟合一个顶层标签BCE温度,将TvA-TS恢复为其K=1极限。在微调的CIFAR-100 / ViT-B/16上,SRTS-BCE将ECE_15从1.65(标量TvA-TS)降至0.96,在全校准预算下匹配更高容量的SMART+BCE头(0.95)。随着预算缩小,两种机制分离:在n=250时,SRTS-BCE在所有三个CIFAR-100骨干网络上优于SMART+BCE(种子到抽样的分层区间排除零),而与标量的旗舰比较仍具方向性。一项协议冻结的Tiny-ImageNet后续实验重现了小预算分离,并在Swin-T上表现出依赖预算的排名反转;匹配的路由和映射控制表明,该效应既与学习路由器无关,也与离散分组无关。综合来看,这些结果将事后校准器容量识别为有限样本设计选择,其首选水平随可用校准数据量而变化。
英文摘要
When a classifier is recalibrated from only a few thousand held-out examples, the capacity of the calibration map becomes a statistical design choice rather than a purely architectural one: a scalar map can underfit structured residual miscalibration, while a highly adaptive map can be hard to estimate reliably from so small a split. We disentangle the calibration objective from adaptive capacity and propose signal-routed temperature scaling (SRTS-BCE), a 10-parameter, argmax-preserving calibrator that cross-fits a correctness-risk score over six logit statistics and fits one top-label-BCE temperature per $K=3$ risk groups, recovering TvA-TS as its $K=1$ limit. On fine-tuned CIFAR-100 / ViT-B/16, SRTS-BCE reduces $\mathrm{ECE}_{15}$ from 1.65 (scalar TvA-TS) to 0.96, matching the higher-capacity SMART+BCE head (0.95) at the full calibration budget. The two regimes separate as the budget shrinks: at $n=250$ SRTS-BCE beats SMART+BCE on all three CIFAR-100 backbones (the seed-to-draw hierarchical interval excludes zero), whereas the flagship comparison against the scalar remains directional. A protocol-frozen Tiny-ImageNet follow-up reproduces the small-budget separation and exhibits a budget-dependent ranking reversal on Swin-T; matched routing and map controls show that the effect is tied neither to the learned router nor to discrete grouping. Together the results identify post-hoc calibrator capacity as a finite-sample design choice whose preferred level shifts with the amount of available calibration data.
发表机构
- Adelaide University(阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。