面向不平衡高风险决策支持的代价敏感共形预测与人在环弃权:多领域基准
Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark
浏览论文内容
中文总结 AI 辅助
该研究针对高风险决策系统的类别不平衡与非对称代价问题,通过多领域基准测试对比多种共形预测及弃权机制,发现Mondrian CP可显著提升少数类覆盖,结合代价控制弃权能降低预期决策代价,为相关系统部署提供实用指导。
中文摘要 AI 辅助
信贷评分、欺诈检测、医疗保健和工业安全等领域的高风险决策系统,需要在严重类别不平衡和非对称误差代价下提供可靠的不确定性量化。标准边际共形预测(CP)可提供有效的整体覆盖保证,但我们发现其对稀有、高代价的少数类覆盖严重不足,在某些数据集上少数类覆盖率低至0.5%。为表征并解决这一局限,我们开展了全面基准测试,在15个现实世界不平衡表格数据集、7种分类模型、3种概率校准技术和10个随机种子上,对比了边际CP、类条件(Mondrian)CP和代价控制弃权机制,共完成3150次实验运行。结果显示,Mondrian CP可恢复有效的少数类覆盖,相比边际CP平均提升61.7个百分点的少数类覆盖率(p < 1e-80)。此外,结合Mondrian CP与代价控制弃权机制,在实际人工审核预算下,相比标准决策边界、基于置信度的拒绝器和风险控制拒绝器,可显著降低预期决策代价。我们还量化了特定数据集的盈亏平衡点阈值,即把模糊实例交由人类专家处理变得具有成本效益的临界点。这些发现为在高风险决策支持系统中部署与分布无关、感知代价的不确定性量化提供了实用指导。
英文摘要
High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs. Standard marginal conformal prediction (CP) provides valid overall coverage guarantees; however, we show that it severely under-covers rare, costly minority classes, with minority-class coverage dropping to as low as 0.5% on certain datasets. To characterize and address this limitation, we conduct a comprehensive benchmark comparing marginal CP, class-conditional (Mondrian) CP, and cost-controlled abstention mechanisms across 15 real-world imbalanced tabular datasets, 7 classification models, 3 probability calibration techniques, and 10 random seeds, resulting in 3,150 experimental runs. Our results show that Mondrian CP restores valid minority-class coverage, achieving an average minority-coverage improvement of 61.7 percentage points over marginal CP (p < 1e-80). Furthermore, combining Mondrian CP with cost-controlled abstention significantly reduces expected decision cost compared with standard decision boundaries, confidence-based rejectors, and risk-controlled rejectors under realistic human review budgets. We further quantify dataset-specific break-even thresholds at which deferring ambiguous instances to human experts becomes cost-effective. These findings provide practical guidance for deploying distribution-free, cost-aware uncertainty quantification in high-stakes decision support systems.
发表机构
- Boston University(波士顿大学)
- University of California(加利福尼亚大学)
- University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。