arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用表格基础模型和System One模型估计未编码的碰撞因素:Kumo Tabular与Jev

Estimating Uncoded Crash Factors with Tabular Foundation and System One Models: Kumo Tabular and Jev

Amir Rafe, Subasish Das

arXiv 2610.10321首次发表:更新:

发表机构

Texas State University(德克萨斯州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究结合Kumo Tabular表格基础模型和Jev System One模型,对德克萨斯州560万起碰撞的编码字段与叙述文本进行联合估计,识别未编码因素并生成验证过的重新阅读列表,提高计数准确性。

AI 中文摘要

道路安全项目统计警察碰撞记录中的编码字段,而警官的叙述(通常记录了字段遗漏的因素)很少被阅读。因此,安全办公室无法知道其统计遗漏了多少,也无法知道应在何处进行审查。本研究开发并评估了一个系统,该系统将2017年至2025年间德克萨斯州5,601,890起碰撞的两种视角结合起来,形成具有声明有效性的总体估计。一个上下文内表格基础模型Kumo Tabular读取每起碰撞的编码记录,一个校准的System One模型Jev读取两个概率样本的叙述,人工判断重新校准其概率。一个多波次“先预测后去偏”估计器将三个层级结合起来,第二个以记录概率抽取的人工层级按设计检查估计值。对于打滑、医疗事件、疲劳、动物和手机使用,叙述记录的伤害碰撞多于编码字段,手机使用为15,074对7,340,人工检查在误差范围内同意所有十五项估计。按Kumo Tabular排序的重新阅读列表发现确认的不一致频率是随机阅读的7至58倍。在人工编码的规划成本下,再增加一轮人工判断将使均方根相对半宽度从22.0%降至16.2%,而阅读所有叙述则为21.2%。两个校准的不同视角阅读器,通过抽样设计结合,为安全办公室提供统计数字、不一致地图、验证过的重新阅读列表和阅读预算,Kumo Tabular读取表格的速度是TabPFN 3.5的15倍。

英文摘要

Road safety programs count the coded fields of police crash records, while the officer's narrative, which often records factors the fields omit, is rarely read. A safety office thus cannot tell how much its counts miss or where to review. This study develops and evaluates a system that joins both views of the 5,601,890 Texas crashes from 2017 to 2025 into population estimates with stated validity. An in-context tabular foundation model, Kumo Tabular, reads the coded record of every crash, a calibrated System One model, Jev, reads the narratives of two probability samples, and human judgments recalibrate its probabilities. A multiwave predict-then-debias estimator joins the three tiers, and a second human tier drawn with recorded probabilities checks the estimates by design. For hydroplaning, medical episodes, fatigue, animals, and phone use, the narrative documents more injury crashes than the coded field, 15,074 against 7,340 for phone use, and the human check agrees with all fifteen estimates within its margin. A re-read list ranked by Kumo Tabular finds confirmed discordance 7 to 58 times as often as random reading. At the planning cost of human coding, one further round of human judgments would cut the root mean square relative half-width from 22.0 to 16.2 percent, against 21.2 for reading every narrative. Two calibrated readers of different views, joined by a sampling design, give a safety office counts, a discordance map, a validated re-read list, and a reading budget, with Kumo Tabular reading the table at 15 times the speed of TabPFN 3.5.

Comments26 pages, 9 figures, 8 tables. Code: https://github.com/pozapas/kumo-jev-crash-records

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑