arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13100stat.MEcs.LGstat.ML

一种用于衡量校准的排序方法

A Ranking Approach for Measuring Calibration

发表机构波士顿大学 · 芝加哥大学
查看机构详情
  • Boston University(波士顿大学)
  • University of Chicago(芝加哥大学)

机构由 AI 辅助整理,请以论文原文为准。

Anirban Chatterjee, Rina Foygel Barber

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种基于排序的校准误差度量rankECE,通过比较预测概率邻近值点,在无假设设定下提供比常用分箱近似更优的ECE代理指标,并给出理论保证与实证验证。

中文摘要 AI 辅助

当使用预测模型提供预测概率时,理想模型应具有完美的校准性:结果的真实概率(即$Y=1$的概率)与预测概率$f(X)$完全匹配。在实践中,模型不可避免地存在校准误差,因此能够衡量这种失准以评估模型的可靠性显得尤为重要。期望校准误差(ECE)是最广泛使用的失准度量,但已知在无假设设定下无法以有保证的精度估计ECE。在本工作中,我们提出了一种替代度量——rankECE,它基于比较预测概率$f(X)$的邻近值点。我们的理论保证和实证结果表明,与实践中常用的分箱近似ECE相比,rankECE提供了更好的ECE代理指标。

英文摘要

When providing forecasted probabilities with a predictive model, the ideal model offers perfect calibration: the true probability of the outcome (i.e., the probability that $Y=1$) exactly matches the forecasted probability $f(X)$. In practice, models inevitably exhibit calibration error, and it is therefore important to be able to measure this miscalibration to assess a model's reliability. The Expected Calibration Error (ECE) is the most widely used measure of miscalibration, but is known to be impossible to estimate the ECE with guaranteed accuracy in an assumption-free setting. In this work, we propose an alternative measure, the rankECE, that is based on comparing points with neighboring values of the predicted probability $f(X)$. Our theoretical guarantees and empirical results establish that rankECE provides a better proxy for ECE as compared to binned approximations to ECE, which are the most commonly-used approximations in practice.

↑