arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向可审计与校准的痴呆相关碰撞严重性预测AI:支持人工审查的选择性延迟框架

Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review

Gaurab Chhetri, Anika Baitullah, Subasish Das

arXiv 2609.22694首次发表:更新:

发表机构

Texas State University(德克萨斯州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种选择性延迟框架,将痴呆相关碰撞严重性预测重构为决策感知分诊问题,在控制泄漏和校准置信度的同时,通过延迟高风险案例支持人工审查,提升自动分类性能与可审计性。

AI 中文摘要

公共碰撞数据库日益支持自动化安全分析,但当模型主要作为普通分类器进行评估时,碰撞严重性预测仍难以转化为公共部门的决策工作流程。本研究将痴呆相关碰撞严重性建模重新定义为一种决策感知的分诊问题,在该问题中,系统必须将碰撞分类为无伤害/仅财产损失(O)、轻微或中度伤害(BC)以及致命或严重伤害(KA),同时控制结果泄漏、报告严重欠分诊、校准置信度,并保留每次原始预测以供审计。利用4,781条德克萨斯州碰撞记录(包含结构化字段和警方叙述),我们在分层70/15/15划分下评估了结构化、叙述、融合、校准融合、BERT系列和本地大型语言模型基线。在报告的划分中,泄漏控制的Gemma获得了最高的观测宏F1分数(0.545;95%自助法置信区间[0.507, 0.583])。最佳校准融合模型获得0.522的宏F1分数和0.033的期望校准误差。选择性延迟提高了保留用于自动分类的案例的性能。在70%覆盖率下,宏F1分数升至0.573,严重性成本降至0.577,而延迟案例被视为拟议人工审查流程的候选对象,在本次实验中未进一步评估。该研究为碰撞AI系统提供了一个可复现、泄漏控制和不确定性感知的评估框架,强调可审计性和选择性延迟而非仅关注准确性。

英文摘要

Public crash databases increasingly support automated safety analysis, but crash severity prediction remains difficult to translate into public-sector decision workflows when models are evaluated primarily as ordinary classifiers. This study reframes dementia-related crash severity modeling as a decision-aware triage problem in which a system must classify crashes into no-injury/property-damage-only (O), minor or moderate injury (BC), and fatal or severe injury (KA), while also controlling outcome leakage, reporting severe under-triage, calibrating confidence, and preserving every raw prediction for audit. Using 4,781 Texas crash records with structured fields and police narratives, we evaluate structured, narrative, fusion, calibrated fusion, BERT-family, and local large-language-model baselines under a stratified 70/15/15 split. In the reported split, leakage-controlled Gemma obtains the highest observed macro-F1 (0.545; 95% bootstrap CI [0.507, 0.583]). The best calibrated fusion model obtains macro-F1 of 0.522 and expected calibration error of 0.033. Selective deferral improves performance among cases retained for automatic classification. At 70% coverage, macro-F1 rises to 0.573 and severity cost falls to 0.577, while deferred cases are treated as candidates for a proposed human-review process and are not further evaluated in the present experiment. The study contributes a reproducible, leakage-controlled, and uncertainty-aware evaluation framework for crash AI systems, emphasizing auditability and selective deferral rather than accuracy alone.

CommentsThis is the author's preprint version of a paper accepted for presentation at HICSS 60 (Hawaii International Conference on System Sciences), 2027, Hawaii, USA. The final published version will appear in the official conference proceedings. Conference site: https://hicss.hawaii.edu/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑