将大型语言模型碰撞严重性流程适配到田纳西州:不同采样策略下的性能表现
Adapting a Large Language Model Crash-Severity Pipeline to Tennessee: Performance Across Sampling Strategies
浏览论文内容
中文总结 AI 辅助
本研究将LLM碰撞严重性工作流适配至田纳西州624,392起碰撞数据,比较随机、基于县和严重性平衡采样策略,发现平衡采样能更均匀识别稀有致命碰撞,强调需结合采样设计与类别指标解读性能。
中文摘要 AI 辅助
各州碰撞数据库在结构、编码和伤害严重性分布上存在差异,限制了预测工作流在不同司法管辖区间的直接复用。本研究将SafeTraffic Copilot大型语言模型(LLM)碰撞严重性工作流适配到包含624,392起碰撞的田纳西州三年期数据库。田纳西州的碰撞、道路、车辆和人员属性被统一并转换为文本提示,缺失值被保留而非推断。使用低秩适配对Llama 3.1 8B进行微调,以分类五种伤害严重性类别。通过独立的样本内和未见测试实验评估了随机、基于县和严重性平衡采样策略;未见测试实验使用70/15/15的训练、验证和测试划分。在各自的未见测试集上,随机和基于县的模型加权F1分数接近80%,但宏F1仍低于47%,致命碰撞F1低于27%,表明强聚合性能可能掩盖对稀有结果的弱识别。在其平衡评估群体中,严重性平衡模型实现了57.7%的加权和宏F1分数,致命碰撞F1分数为65.9%,产生了更均匀的类别级性能。由于每种采样策略使用了不同的测试子集,跨策略差异是描述性的而非受控排名。结果强调了将LLM碰撞严重性性能与采样设计、类别平衡、类别级指标和评估群体构成一起解读的重要性。
英文摘要
State crash databases differ in structure, coding, and injury-severity distributions, limiting direct reuse of predictive workflows across jurisdictions. This study adapts the SafeTraffic Copilot large language model (LLM) crash-severity workflow to a three-year Tennessee inventory of 624,392 crashes. Tennessee crash, roadway, vehicle, and person attributes were harmonized and converted into textual prompts while unavailable values were preserved rather than inferred. Llama 3.1 8B was fine-tuned using low-rank adaptation to classify five injury-severity categories. Random, county-based, and severity-balanced sampling strategies were evaluated using separate in-sample and unseen-test experiments; unseen-test experiments used 70/15/15 training, validation, and test splits. On their respective unseen test sets, random and county-based models achieved weighted F1-scores near 80%, but macro F1 remained below 47% and fatal-crash F1 below 27%, showing that strong aggregate performance can mask weak recognition of rare outcomes. Within its balanced evaluation population, the severity-balanced model achieved weighted and macro F1-scores of 57.7% and a fatalcrash F1-score of 65.9%, yielding more even class-level performance. Because each sampling strategy used a different test subset, crossstrategy differences are descriptive rather than controlled rankings. The results highlight the importance of interpreting LLM crashseverity performance together with sampling design, class balance, class-level metrics, and evaluation-population composition.
发表机构
- Oak Ridge National Laboratory(橡树岭国家实验室)
- University of Tennessee at Knoxville(田纳西大学诺克斯维尔分校)
- University of Tennessee - Oak Ridge Innovation Institute(田纳西大学-橡树岭创新学院)
机构由 AI 辅助整理,请以论文原文为准。