模型信息不足下可靠性集估计中的可信决策
Trustworthy Decisions in Reliability Set Estimation under Insufficient Model Information
浏览论文内容
中文总结 AI 辅助
本研究针对模型信息不足的可靠性集估计问题,提出统一框架,通过建模后校准、共形风险控制、自适应设计实现可信决策,获更准确集估计与校准风险控制。
中文摘要 AI 辅助
可靠性集估计用于识别响应概率超过目标水平的输入区域,是估计与安全关键决策之间的桥梁。从业者通常从工作模型(真实响应曲面的不完美近似)出发,依赖该不完美模型可能产生决策风险,或将不安全区域误判为安全。我们开发了一个统一框架,将此类工作模型转化为可信决策规则。首先,采用“建模后校准”程序将估计与决策解耦;由于真实集不可观测,我们引入非对称、可观测的代理损失,并使用单独的校准集选择偏差校正阈值,降低决策风险,实现$O_P(1/n)$的体积收敛。其次,利用基于代理损失的共形风险控制,将最具安全关键性的错误——错误包含风险控制在预先指定的水平,且不受工作模型质量影响。这些校准程序共同表明,单独的校准集对风险控制是必要的。第三,自适应设计将观测集中在可靠性集及其边界,在错误对决策影响最大的区域提升模型质量,同时控制其他区域的预算。数值研究显示,与经典插入法相比,该方法不仅能得到更准确的集估计,还具备有限样本下的校准风险控制能力。
英文摘要
Reliability set estimation identifies input regions where a response probability exceeds a target level, bridging estimation and safety-critical decisions. Practitioners typically start with a working model, an imperfect approximation of the true response surface. Relying on this imperfect model may incur decision risk, potentially certifying unsafe regions as safe. We develop a unified framework that turns such a working model into a trustworthy decision rule. First, a modeling-then-calibration procedure decouples estimation from decision. Since the true set is unobservable, we introduce an asymmetric, observable surrogate loss and use a separate calibration set to select a bias-correcting threshold, reducing decision risk and achieving $O_P(1/n)$ volume convergence. Second, we leverage conformal risk control with the surrogate loss to control false inclusion risk, which is the most safety-critical error, at a pre-specified level regardless of working model quality. Together, these calibration procedures show that a separate calibration set is necessary for risk control. Third, an adaptive design concentrates observations on the reliability set and its boundary, improving model quality where errors most affect decisions while controlling budget elsewhere. Numerical studies show not only more accurate set estimates but also calibrated finite-sample risk control that classical plug-in methods lack.