arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

成本感知的后验延迟决策在标定与偏移下的研究:一个环境人工智能案例

Cost-Aware Post-Hoc Deferral Under Calibration and Shift: An Environmental AI Case Study

Haoran Yu, Lifei Liu, Danping Zhang

arXiv 2609.09235首次发表:更新:

发表机构

University of Florida; Wichita State University; Nanchang Hangkong University(佛罗里达大学; 威奇托州立大学; 南昌航空大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出EcoTrust框架,结合风险估计与置信度规则,在环境AI案例中比较自动操作与审查的成本,发现丰富风险信号未必优于标定置信度。

AI 中文摘要

为冻结的分类器选择延迟决策策略,不仅需要对不确定案例进行排序:置信度可能未被正确标定,错误具有不等的成本,审查者可能出错,且部署数据可能超出标定支持范围。我们通过EcoTrust框架研究这些交互作用,该框架是一个后验框架,使用六组错误风险估计器、类别不对称成本、审查者准确性以及可选的支持门控,将自动操作与审查进行比较。在哥伦比亚河热应力测试平台上,学习到的估计器将错误排序的接受者操作特征曲线下面积从0.869提高到0.889,但Chow的置信度规则具有更低的分布内成本(每天0.416对比0.567)。在12个现成后端中,学习到的风险以及标定匹配的、类别感知的置信度估计器在六个后端上各自优于原始Chow;配对的年份块自助法无法解决其平均成本差异。在转移到十个河流站点时,门控标记每个案例并成为始终审查的回退机制,仅在审查完美且不受限制时,在八个站点上达到最低成本。这些结果刻画了一个受控任务上的决策边界:更丰富的风险信号并不能可靠地优于标定后的置信度,且检测到的外推并不暗示可转移的案例级排序。

英文摘要

Choosing a deferral policy for a frozen classifier requires more than ranking uncertain cases: confidence may be miscalibrated, errors have unequal costs, reviewers can err, and deployment data can leave calibration support. We study these interactions through EcoTrust, a post-hoc framework that compares automatic action with review using a six-group error-risk estimator, class-asymmetric costs, reviewer accuracy, and an optional support gate. On a Columbia River thermal-stress testbed, the learned estimator improves error-ranking area under the receiver operating characteristic curve from 0.869 to 0.889, but Chow's confidence rule has lower in-distribution cost (0.416 versus 0.567 per day). Across 12 off-the-shelf backends, learned risk and a calibration-matched, class-aware confidence estimator each beat raw Chow on six; a paired year-block bootstrap does not resolve their mean cost difference. In transfer to ten river stations, the gate flags every case and becomes an always-review fallback, attaining the lowest cost on eight stations only when review is perfect and unconstrained. These results characterize decision boundaries on one controlled task: richer risk signals do not reliably improve on calibrated confidence, and detected extrapolation does not imply transferable case-level ranking.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑