arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24052cs.CL

大规模校准决策:使用System One模型(Jev)将警方碰撞叙述转化为概率性碰撞变量

Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev)

  • Ingram School of Engineering, Texas State University(德克萨斯州立大学英格拉姆工程学院)

机构由 AI 辅助整理,请以论文原文为准。

Amir Rafe, Subasish Das

AI总结:

本文提出用System One模型Jev将警方碰撞叙述编码为概率变量,在49.95万条德克萨斯叙述上实现F1=0.908,校准误差降低3.3倍,并证明该方法可扩展且成本可控。

AI中文摘要:

带有调查员叙述的碰撞数据集包含了编码字段所遗漏的信息。大规模地对这些叙述进行编码一直受到三个障碍的阻碍。前沿大型语言模型在该规模下成本高昂,其生成的文本无法验证,并且没有规则规定人类必须检查多少输出。本文将叙述编码表述为门控的、类型化的决策,由Jev(一种System One模型)回答,该模型返回分析师定义选项上的概率,且不生成文本。一项筛查覆盖了499,500条德克萨斯州叙述,其中195,857条使用27个问题的模式进行了编码。成本由模式大小而非叙述长度决定。这些概率对照编码字段以及根据既定抽样设计得出的2,416个盲法人工判断进行了审计。两个前沿大型语言模型在相同记录上进行了基准测试。针对人工标签,该类型化模型的F1分数达到0.908。一个前沿模型提高了0.059,另一个与其无法区分。校准效果因模型而异,而非因范式而异,因此每个模型都必须经过审计。对相同标签进行重新校准可将校准误差降低3.3倍。与编码字段的一致性低估了叙述保真度,kappa值中位数差异为0.26。一个分辨率下限界适用于任何在离散网格上报告概率的模型。对标记记录的审查预算给出了每个变量每年人类必须阅读的记录数。将校准变量添加到编码字段中,使归因于九个因素的伤害和致命碰撞每年增加10,747起。

英文摘要:

Crash datasets that carry an investigator narrative hold information the coded fields omit. Coding those narratives at scale has been blocked by three obstacles. Frontier large language models are costly at that scale, their generated text cannot be verified, and no rule says how much output a human must check. This paper formulates narrative coding as gated, typed decisions answered by Jev, a System One model that returns probabilities over analyst-defined options and generates no text. A screen covered 499,500 Texas narratives and 195,857 were coded with a 27-question schema. Cost is governed by schema size rather than narrative length. The probabilities are audited against coded fields and against 2,416 blinded human judgments drawn under a stated sampling design. Two frontier large language models are benchmarked on the same records. Against human labels the typed model attains an F1 of 0.908. One frontier model gains 0.059 and the other is indistinguishable from it. Calibration varies by model rather than by paradigm, so each model must be audited. Recalibration on the same labels reduces calibration error by a factor of 3.3. Agreement with coded fields understates fidelity to the narrative by a median of 0.26 in kappa. A resolution-floor bound covers any model that reports probabilities on a discrete grid. A review budget over flagged records gives the records a human must read per variable and per year. Adding the calibrated variables to the coded fields raises the injury and fatal crashes attributed to nine factors by 10,747 per year.

↑