发表机构
University of South Florida(南佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出VLAlert框架,将驾驶员告警建模为静默、观察、告警三动作策略,利用视觉语言模型生成安全证据并自适应观察,在多个基准上显著提升告警效用与性能。
AI 中文摘要
从行车记录仪视频进行驾驶员告警需要在部分可观测条件下进行序贯决策:系统不仅要判断场景是否危险,还要判断何时证据充分足以发出警告。现有的大多数事故预测模型输出二元风险分数,对模糊场景的处理依赖于阈值设定。我们提出VLAlert,一个视觉语言告警框架,将警告生成建模为SILENT(静默)、OBSERVE(观察)和ALERT(告警)三动作策略。OBSERVE动作作为内部证据收集决策,延迟不确定的警告并改变下一个观察窗口,从而构建一个轻量级的感知-行动循环以实现自适应告警。VLAlert使用Qwen3-VL-4B作为安全证据生成器,并从结构化信念跨度中池化隐藏状态,形成用于危险估计和策略预测的紧凑表示。我们在VLAlert-Bench上评估VLAlert,这是一个基于四个真实行车记录仪告警数据集构建的统一逐帧基准,并进一步测试其对保留的自然驾驶ADAS接管片段的迁移能力。在VLAlert-Bench验证集上,VLAlert在测试基线中取得了最高的部署导向效用,DAUS为0.4878,而Open-BADAS为0.4752,并将AUROC、AP_tick、F1_t和平衡准确率分别从0.610、0.176、0.276和0.581提升至0.689、0.195、0.297和0.648。在221个保留的ADAS-TO-Critic片段上,VLAlert将R@5s从74.2%提升至88.7%,F1从0.585提升至0.686。这些结果表明,自适应观察和以安全为中心的VLM表示为面向驾驶员的告警决策带来了可衡量的改进。
英文摘要
Driver alerting from dashcam video requires sequential decision-making under partial observability: a system must decide not only whether a scene is risky, but also when the evidence is sufficient to warn. Most existing accident anticipation models output a binary risk score, leaving ambiguous scenes to be handled by thresholding. We propose VLAlert, a vision-language alerting framework that casts warning generation as a tri-action policy over SILENT, OBSERVE, and ALERT. The OBSERVE action acts as an internal evidence-gathering decision that delays uncertain warnings and changes the next observation window, creating a lightweight perception-action loop for adaptive alerting. VLAlert uses Qwen3-VL-4B as a safety-evidence generator and pools hidden states from structured belief spans to form compact representations for danger estimation and policy prediction. We evaluate VLAlert on VLAlert-Bench, a unified per-tick benchmark from four real-world dashcam alert datasets, and further test transfer to held-out naturalistic ADAS takeover clips. On VLAlert-Bench validation, VLAlert achieves the highest deployment-oriented utility among tested baselines, with DAUS 0.4878 compared with 0.4752 for Open-BADAS, and improves AUROC, AP_tick, F1_t, and balanced accuracy from 0.610, 0.176, 0.276, and 0.581 to 0.689, 0.195, 0.297, and 0.648, respectively. On 221 held-out ADAS-TO-Critic clips, VLAlert improves R@5s from 74.2% to 88.7% and F1 from 0.585 to 0.686. These results indicate that adaptive observation and safety-focused VLM representations provide measurable gains for driver-facing alert decisions.
Comments23 pages, 8 figures. Accepted at the Conference on Robot Learning (CoRL) 2026