发表机构
Julius-Maximilians-Universität Würzburg; Zaragoza Logistics Center; INSEAD(维尔茨堡大学朱利叶斯 - 马克西米利安斯大学; 萨拉戈萨物流中心; 欧洲工商管理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对黑箱条件分位数预测器,开发无分布且基于博弈论的测试框架,通过形式化校准概念、识别可实现功效的替代方案集,得出有限时间检测保证,能在特征层面解释证据过程,实证发现流行预测器存在校准错误。
AI 中文摘要
黑箱条件分位数预测在非对称成本下的顺序决策中广泛应用,如供应链管理中的库存规划。部署后,因数据流漂移和模式变化需持续监控,这使标准固定视野回测校准失效,且现有回测未考虑校准与信息相关。我们开发了无分布且基于博弈论的测试框架,用于持续审计具有非独立同分布损失的黑箱条件分位数预测器。先形式化不同特征集下的条件分位数校准概念,确定审计信息集的粗糙程度决定测试问题的难度,然后识别审计可实现功效的替代方案集,针对基于特征线性的上下文赌注得出有限时间检测保证,且无需独立同分布假设。所得证据过程在特征层面可解释,能量化校准错误的细粒度“特征感知”证据。我们在模拟和真实数据上实证验证了这些方法,发现流行的时间序列预测器(Chronos - 2)在多个相关特征上高度校准错误。
英文摘要
Black-box conditional quantile forecasts are widely used for sequential decisions under asymmetric costs, such as inventory planning in supply chain management. Once deployed, such forecasters must be monitored continuously as data streams drift and regimes change; this invalidates standard, fixed-horizon backtests for calibration. Further, existing backtests do not take into account that the notion of calibration is, in fact, information-dependent: forecasts can look calibrated to an auditor with coarse information while being miscalibrated to an auditor with richer information. We develop a distribution-free and game-theoretic testing framework for continuously auditing black-box conditional quantile forecasters with non-i.i.d. losses, such that the resulting evidence process is powerful against predictably chosen alternatives specified by the features available to the auditor. We first formalize notions of conditional quantile calibration when different sets of features are available to the auditor, establishing that the coarseness of the auditor's information set determines the hardness of the testing problem. We then identify the sets of alternatives for which the auditor can achieve power, and focusing on contextual bets linear in the features, we derive finite-time detection guarantees for such alternatives, all without an i.i.d. assumption. The resulting evidence processes are interpretable at the feature level, as they quantify fine-grained, "feature-aware" evidence for miscalibration. We empirically validate these methods on simulated and real data, finding that a popular time series forecaster (Chronos-2) is highly miscalibrated w.r.t. multiple relevant features.