arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自适应日志下的反事实在线共形预测

Counterfactual Online Conformal Prediction Under Adaptive Logging

Xinyu Qiao, Yichen Lin, Kaihong Ji, Xue Wang, Tao Yao

arXiv 2609.30811首次发表:更新:

发表机构

Shanghai Jiao Tong University; Alibaba Group(上海交通大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出倾向加权在线共形预测(PW-OCP)及其双稳健变体(DR-OCP),以解决自适应日志下反事实覆盖失效问题,实验证明其提升反事实覆盖并降低遗憾。

AI 中文摘要

当预测影响行动,而行动决定哪些结果进入校准集时,在线共形预测可能失效。标准的自适应方法可能保持边际覆盖,却系统性地漏覆盖那些极少被选择行动的反事实结果。本文通过反事实覆盖形式化这一失效,并引入倾向加权在线共形预测(Propensity-Weighted Online Conformal Prediction, PW-OCP),一种逆倾向加权递归,用于去偏校准。一种双稳健变体进一步将干扰偏差降低至结果模型与倾向性误差的乘积。在正性假设下,所得覆盖率匹配信息论下界(相差对数因子)。在合成决策任务、开放bandit数据和金融再平衡上的实验表明,PW-OCP和DR-OCP在不牺牲预测集锐度的情况下,改善了反事实覆盖率和下游遗憾。

英文摘要

Online conformal prediction can fail when predictions shape actions and actions determine which outcomes enter calibration. Standard adaptive methods may retain marginal coverage while systematically miscovering the counterfactual outcomes of rarely selected actions. This paper formalizes the failure through counterfactual coverage and introduces Propensity-Weighted Online Conformal Prediction, an inverse-propensity-weighted recursion that debiases calibration. A doubly robust variant further reduces nuisance bias to the product of outcome-model and propensity errors. Under positivity, the resulting coverage rate matches an information-theoretic lower bound up to logarithmic factors. Experiments on synthetic decision tasks, open bandit data, and financial rebalancing show that PW-OCP and DR-OCP improve counterfactual coverage and downstream regret without sacrificing prediction-set sharpness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑