发表机构
Binghamton University (SUNY); School of Computing(纽约州立大学宾汉姆顿分校; 计算机学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究软标签贝叶斯误差估计器在概率非真实后验时的脆弱性,通过刻画温度缩放对代理的影响,证明相关恒等式并推导封闭形式,在多数据集上验证,给出失真预测,强调代理值与概率产生机制结合才有意义。
AI 中文摘要
石田等人的软标签贝叶斯误差估计器β(z)=E[min(z,1 - z)]直接从概率值标签估计二元任务的不可约误差。Ushio等人最近的工作表明,当概率不是真实后验时,该估计器很脆弱。我们通过精确刻画最常用的事后校准映射——温度缩放——如何扭曲代理来补充这项工作。我们证明了一个精确的、无模型的恒等式,将温度缩放后的代理简化为分类器的边缘分布,由此我们得到了温度的严格单调性以及从温度轴到开区间(0,1/2)的连续双射。在对数its的高斯模型下,我们进一步推导出整个代理与温度曲线的双参数封闭形式。在CIFAR - 10、Fashion - MNIST和SVHN(八个二元任务)上,代理在恒定测试误差下变化56倍到980倍,封闭形式将经验曲线的再现误差控制在0.018以内,并且使预期校准误差最小化的校准温度与任何稳定的代理值都不一致。我们的结果给出了失真的精确预测描述,强化了只有代理值与其产生概率的机制一起才有意义的实际建议。
英文摘要
The soft-label Bayes-error estimator beta(z) = E[min(z, 1-z)] of Ishida et al. estimates the irreducible error of a binary task directly from probability-valued labels. Recent work by Ushio et al. showed that this estimator is fragile when the probabilities are not the true posterior: even perfectly calibrated soft labels can yield a substantially inaccurate estimate, and they propose isotonic calibration as a consistent remedy. We complement that line of work by characterizing exactly how the most widely used post-hoc calibration map -- temperature scaling -- distorts the proxy. We prove an exact, model-free identity reducing the temperature-scaled proxy to the classifier's margin distribution, from which we obtain (i) strict monotonicity in the temperature and (ii) a continuous bijection from the temperature axis onto the open interval (0, 1/2), so that a fixed classifier -- with fixed decisions and fixed 0-1 error -- can be made to report any proxy value whatsoever. Under a Gaussian model of the logits we further derive a two-parameter closed form for the entire proxy-versus-temperature curve. Across CIFAR-10, Fashion-MNIST, and SVHN (eight binary tasks), the proxy varies by 56x to 980x at constant test error, the closed form reproduces the empirical curve to within 0.018, and the calibration temperature that minimizes the expected calibration error does not coincide with any stable proxy value. Our results give a precise, predictive account of the distortion whose existence motivates calibration-based remedies, and they reinforce the practical recommendation that a proxy value is meaningful only together with the mechanism that produced its probabilities.
Comments6 pages, 2 figures. Code available at https://github.com/sherurox/bayes-proxy-channel