arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于警方事故叙述文本,对比前沿大语言模型与官方事故数据库编码的基准测试

Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives

Sudhir Bharati, Rajendra K C Khatri, Sudip Bharati

arXiv 2607.29064首次发表:更新:

AI 中文总结

本研究以阿肯色州2015-2025年的致命事故数据为对象,对比6种前沿大语言模型与官方事故数据库的编码效果,为事故编码领域的LLM应用提供了基准评估依据。

AI 中文摘要

警方事故叙述文本包含可补充结构化事故数据库的信息,但人工审查耗时费力,且大语言模型(LLMs)重现官方事故编码的效果尚不明确。本研究以阿肯色州2015-2025年的致命事故为研究对象,将5587份致命事故叙述文本与5889条结构化事故记录关联,得到4194起匹配事故,通过对比由叙述文本推导的事故属性编码与官方数据库对应字段,对6种前沿LLMs开展基准测试。所有LLMs采用相同的零样本提示词,编码事故类型、非机动车相关情况、路口类型、施工区域相关情况、路面状况及光照条件,采用一致性、宏平均F1值、Cohen's kappa、覆盖率、选择性一致性等指标,结合始终选多数类、始终选未知类、关键词规则基线进行评估,同时用重复测量分析和广义估计方程模型分析模型间及属性间的差异。结果显示,GPT-5.5 High在评估的LLMs中一致性最高,但始终选多数类基线的原始一致性更高,关键词规则基线的宏平均F1值和Cohen's kappa与表现最优的LLMs相当;非机动车相关情况和事故类型的一致性最高,光照条件、路面状况及施工区域相关情况的一致性最低,事故属性间的差异超过模型间的差异。这些结果为评估基于LLM的事故编码提供了基准,表明部署此类工具需基于属性特异性,结合透明基线和人工审查开展评估。

英文摘要

Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it remains unclear how well large language models (LLMs) reproduce official crash coding. This study benchmarked six frontier LLMs by comparing narrative-derived crash attribute codes with corresponding fields in the Arkansas fatal-crash database. The analysis linked 5,587 fatal-crash narratives with 5,889 structured crash records from Arkansas (2015-2025), yielding 4,194 matched crashes. Six LLMs were evaluated using an identical zero-shot prompt to code crash manner, non-motorist relation, intersection type, work-zone relation, roadway surface condition, and light condition. Performance was evaluated using agreement, macro-averaged F1 score, Cohen's kappa, coverage, selective agreement, and comparisons with always-majority, always-Unknown, and keyword-rule baselines. Repeated-measures analyses and a generalized estimating equations model assessed differences among models and attributes. GPT-5.5 High achieved the highest agreement among the evaluated LLMs, but the always-majority baseline produced higher raw agreement and the keyword-rule baseline achieved macro-averaged F1 score and Cohen's kappa comparable to the best-performing LLM. Agreement was highest for non-motorist relation and crash manner and lowest for light condition, roadway surface condition, and work-zone relation. Differences across crash attributes exceeded differences across models. These results provide a benchmark for evaluating LLM-based crash coding and show that deployment should be evaluated on an attribute-specific basis using transparent baselines and human review.

Comments16 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑