U-Calibration:面向未知智能体的预测
U-Calibration: Forecasting for an Unknown Agent
- Cornell University(康奈尔大学)
- Google Research(谷歌研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对未知效用智能体的预测评估问题,提出U-calibration指标及在线算法,证明其误差是所有智能体实现次线性遗憾的充要条件,且达到最优速率。
AI中文摘要:
我们考虑评估二元事件预测的问题,这些预测被理性智能体消费,智能体会根据预测采取行动,但其效用对预测者来说是未知的。我们表明,针对单一评分规则(例如 Brier score)优化预测,无法保证对所有可能的智能体都具有低遗憾。相比之下,校准良好的预测能保证所有智能体承受次线性遗憾。然而,校准在此并非必要标准(未校准的预测也可能为所有可能的智能体提供良好的遗憾保证),并且已证明校准预测过程的收敛速率比针对单一评分规则的预测过程更差。受此启发,我们提出了一种评估预测的新指标,称为 U-calibration,等于在任何有界评分规则下评估时,预测序列的最大遗憾。我们证明,次线性 U-calibration 误差是所有智能体实现次线性遗憾保证的充要条件。此外,我们展示了如何高效计算 U-calibration 误差,并提供了一种在线算法,实现了 $O(\sqrt{T})$ 的 U-calibration 误差(与针对单一评分规则优化的最优速率相当,并绕过了传统校准学习过程的下界)。最后,我们讨论了向多类预测设置的推广。
英文摘要:
We consider the problem of evaluating forecasts of binary events whose predictions are consumed by rational agents who take an action in response to a prediction, but whose utility is unknown to the forecaster. We show that optimizing forecasts for a single scoring rule (e.g., the Brier score) cannot guarantee low regret for all possible agents. In contrast, forecasts that are well-calibrated guarantee that all agents incur sublinear regret. However, calibration is not a necessary criterion here (it is possible for miscalibrated forecasts to provide good regret guarantees for all possible agents), and calibrated forecasting procedures have provably worse convergence rates than forecasting procedures targeting a single scoring rule. Motivated by this, we present a new metric for evaluating forecasts that we call U-calibration, equal to the maximal regret of the sequence of forecasts when evaluated under any bounded scoring rule. We show that sublinear U-calibration error is a necessary and sufficient condition for all agents to achieve sublinear regret guarantees. We additionally demonstrate how to compute the U-calibration error efficiently and provide an online algorithm that achieves $O(\sqrt{T})$ U-calibration error (on par with optimal rates for optimizing for a single scoring rule, and bypassing lower bounds for the traditionally calibrated learning procedures). Finally, we discuss generalizations to the multiclass prediction setting.