一个用于可信碰撞严重性预测的无分布认证框架
CHOIR: heterogeneity-aware conformal prediction for crash injury severity across driver safety strata
浏览论文内容
中文总结 AI 辅助
针对碰撞严重性预测缺乏有限样本保证的问题,提出一个包装任意模型的无分布认证层,提供序数集、逐类有效性、覆盖转移等保证,并在520万条记录上验证其有效性。
中文摘要 AI 辅助
碰撞严重性模型为筛查、调度和地点优先级排序提供信息,然而这些模型在部署时并未附带关于单次预测含义的有限样本声明。现成的保证在此失效,因为使碰撞严重性具有独特性的特征击败了这些保证:KABCO结果是序数型的,记录标签是一种现场评估,其与医学严重性的一致性约有一半时间,且以结构化方式出错,部署跨越了校准从未见过的司法管辖区和年份。我们开发了一个认证层,它包装任何未经修改的严重性模型,利用这一结构提供无分布保证:连续的序数集合,读作“B或更差”;对任何预先声明的分区提供逐类有效性,并具有最优效率特征;通过声明的报告带将覆盖范围转移到未观测的真实严重性,并具有最坏情况下的锐度结果;在部署偏移下提供单侧证书;以及严重性加权的风险控制。这些保证与可归因的松弛预算组合。同样的分析界定了认证所能达到的极限。认证集合的信息量受真实规律的一个泛函控制,该泛函是任何基础模型都无法规避的,且无法在无分布情况下给出下界;给定一个从记录链接数据中识别出的声明的误报通道,该下限的非空下界变得可计算。在跨越四个十年的七个基础模型上的520万条德克萨斯记录上,该层附加了相同的有效性,并在弱势道路使用者上认证了一个模型无关的集合宽度下限,任何基础模型都无法超越,将其与一个保持有界但无分布不可识别的剩余部分区分开来。该框架作为带有定理级测试的开源软件包发布。
英文摘要
Transportation agencies increasingly predict crash-injury severity with statistical and machine-learning models, but these models do not state how often their output contains the recorded injury level or for which groups of drivers it fails, a gap that matters most for motorcyclists and unrestrained drivers. This study develops and evaluates a certification layer that gives any fitted severity model a finite-sample, distribution-free coverage guarantee within prespecified safety strata. The layer, CHOIR (Conformal Heterogeneity-aware Ordinal Inference with Risk control), combines groupwise and weighted conformal prediction with conformal risk control to return contiguous KABCO intervals, and adds a declared sensitivity analysis for medically assessed injury and bounds on fatal omission. It is evaluated on 4.04 million Texas crashes from 2017-2023, one sampled driver per crash, with seven base models from the ordered logit to a tabular foundation model, and on held-out counties and later years. Under one pooled threshold every model reaches 0.90 coverage overall but covers motorcyclists or unrestrained drivers at 0.868 or lower, and class-balanced gradient boosting covers unrestrained drivers at only 0.374. Calibration within four safety strata places all 28 model-by-stratum estimates between 0.898 and 0.907, at the cost of sets spanning 3.6 to 4.6 of five categories for these groups, and the certified ordered logit is within 0.03 categories of the narrowest model. Injury-model coverage should therefore be certified within safety groups rather than on average, calibration rather than model complexity determines validity, and a statewide threshold should not be applied to small rural counties without local calibration data.
发表机构
- Texas State University(德克萨斯州立大学)
机构由 AI 辅助整理,请以论文原文为准。