arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EG-ARSA:适用于低资源场景的基于专家知识的开放视觉道路安全审计模型

EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings

Md Thamed Bin Zaman Chowdhury, Moazzem Hossain

arXiv 2608.23563首次发表:更新:

发表机构

Bangladesh University of Engineering and Technology (BUET)(孟加拉工程技术大学(BUET))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对中低收入国家道路安全审计的资源限制,提出EGD框架,构建BD-ARSA数据集和EG-ARSA模型,实验显示其性能优于310亿参数教师模型及Gemini-2.5-Flash,为低资源场景道路安全审计提供有效方案。

AI 中文摘要

道路交通伤害仍是中低收入国家面临的重大挑战,这些国家的主动道路安全审计受限于事故记录不完整、合格审计人员短缺以及大规模现场检查成本高昂。为解决该问题,我们提出Expert-Grounded Distillation(EGD),这是一种新型人工智能框架,可将机构道路安全专业知识迁移至紧凑的视觉语言模型,以实现可扩展的视觉道路安全审计。其核心创新在于量化的专家知识对齐阶段,在此阶段,教师视觉语言模型需与权威现场审计结果校准,仅当教师模型与专家风险评估达成高度一致(Cohen's kappa=0.74)时,才允许进行大规模标注。校准后的教师模型随后生成结构化监督信号,通过Low-Rank Adaptation和单个无泄漏提示,将其蒸馏为80亿参数的学生视觉语言模型。我们还推出了Bangladesh Road Safety Audit(BD-ARSA),这是首个基于专家知识的孟加拉国开放视觉道路安全审计数据集,包含近全国覆盖的21947条图像-审计记录;以及Expert-Grounded Road Safety Auditor(EG-ARSA),这是首个专门针对该任务开发的视觉语言模型。实验结果表明,与零样本基线相比,基于知识的微调显著提升了序数风险评估性能,而盲测专家评估显示,该紧凑学生模型的表现优于其310亿参数的教师模型和Gemini-2.5-Flash。这些研究结果表明,EGD为资源受限环境下的主动道路安全审计提供了一种有效且可扩展的工程解决方案。

英文摘要

Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety auditing is limited by incomplete crash records, shortages of qualified auditors, and the high cost of large-scale field inspections. To address this problem, we propose Expert-Grounded Distillation (EGD), a novel artificial intelligence framework that transfers institutional road safety expertise into a compact vision-language model for scalable visual road safety auditing. The key innovation is a quantified expert-grounding stage in which the teacher vision-language model is calibrated against authoritative field audits. Large-scale annotation is permitted only after the teacher reaches substantial agreement with expert risk assessments (Cohen's kappa = 0.74). The calibrated teacher then generates structured supervision that is distilled into an 8-billion-parameter student vision-language model using Low-Rank Adaptation and a single leakage-free prompt. We also introduce Bangladesh Road Safety Audit (BD-ARSA), the first open, expert-grounded Bangladeshi visual road safety audit dataset containing 21,947 image-audit records with near-national coverage, and Expert-Grounded Road Safety Auditor (EG-ARSA), the first vision-language model developed specifically for this task. Experimental results show that grounded fine-tuning substantially improves ordinal risk assessment over the zero-shot baseline, while blind expert evaluation demonstrates that the compact student outperforms both its 31 billion-parameter teacher and Gemini-2.5-Flash. These findings demonstrate that EGD provides an effective and scalable engineering solution for proactive road safety auditing in resource-constrained environments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑