超越反馈的在线共形预测
Online Conformal Prediction Beyond Feedback
浏览论文内容
中文总结 AI 辅助
针对超越反馈的在线共形预测场景,本文提出OCPQ方法,将标签高效预测器适配到该设置,实现了低查询率下的高覆盖率,且在真实数据集上验证了其有效性。
中文摘要 AI 辅助
不确定性量化对于将机器学习模型部署到安全关键型应用中至关重要。在线共形预测(Online Conformal Prediction, OCP)可为任意黑盒分类器和非独立同分布(non-i.i.d.)数据流提供具有理论依据的不确定性量化,其通过构建预测集,确保预测集以用户指定的频率包含真实标签。OCP通常利用先前部署预测的反馈来更新预测集,而我们研究的是超越反馈的OCP设置:在每一轮中,学习器要么输出预测集,要么查询真实标签,但不能同时执行这两种操作,因此没有任何部署的预测会被直接评估。我们将该问题归约为部分监测博弈,其中预测动作不返回观测值,而单独的查询动作会揭示标签。奖励函数的构建方式为:鼓励学习器输出较小的预测集,同时确保真实标签被覆盖的概率足够高。为解决该博弈,我们将Cesa-Bianchi、Lugosi和Stoltz(2004)提出的标签高效预测器适配到我们的设置中,开发了带查询的OCP(OCP with queries, OCPQ)。对于任意黑盒分类器和任意长度为$T$的(non-i.i.d.)遗忘数据流,OCPQ具有$O(T^{2/3})$的期望遗憾,且对于用户定义的$\beta$,其期望覆盖率至少为$\beta - O(T^{-1/3})$,同时仅查询期望占比为$T^{-1/3}$的轮次。这提供了与基于多臂赌博机的OCP方法相当的覆盖率,同时无需从部署的预测集获取反馈。在真实世界数据集上的实验进一步证明了我们方法的有效性。
英文摘要
Uncertainty quantification is essential when deploying machine learning models in safety-critical applications. Online conformal prediction (OCP) provides theoretically principled uncertainty quantification for arbitrary black-box classifiers and non-i.i.d. data streams by constructing prediction sets that are guaranteed to contain the true label at a user-specified frequency. OCP usually updates prediction sets using feedback from previously deployed predictions. We instead study an OCP setting beyond feedback: on each round, the learner can either output a prediction set or query the correct label, but not both. Thus, no deployed prediction is ever evaluated directly. We reduce this problem to a partial monitoring game in which prediction actions return no observation and a separate query action reveals the label. The reward function is constructed in a way that encourages the learner to output small prediction sets while ensuring that the correct label is covered with a sufficiently high probability. To solve this game, we develop OCP with queries (OCPQ) by adapting the label efficient forecaster of Cesa-Bianchi, Lugosi, and Stoltz (2004) to our setting. For any black box classifier and any (non-i.i.d.) oblivious data stream of length $T$, OCPQ has $O(T^{2/3})$ expected regret and expected coverage at least $β-O(T^{-1/3})$ for a user-defined $β$, while querying only an expected $T^{-1/3}$ fraction of rounds. This provides coverage comparable to bandit-based OCP methods while requiring no feedback from deployed prediction sets. Experiments on real-world datasets further demonstrate the effectiveness of our approach.