arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05560cs.CVcs.CL

从运动到安全:多模态大语言模型(MLLM)主动风险推理基准测试

From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs

  • School of Information, Renmin University of China(中国人民大学信息学院)

机构由 AI 辅助整理,请以论文原文为准。

Jiawei Qiu, Yichen Xu, Jianzhe Ma, Mingyang Yu, Wenbin Zhu, Yang Han, Pinzheng Lv, Wenxuan Wang

AI总结:

该研究构建了运动主动风险推理基准SPRINT,评估发现当前MLLM危险感知率高但原因识别率低,仅具备表层主动安全能力,缺乏稳定的基于原因的早期预警。

AI中文摘要:

及时预判物理危险对现实世界安全至关重要,但现有多模态大语言模型(MLLM)评估聚焦于有害内容或一般风险,主动物理危险预测领域仍未得到充分探索。运动场景是理想的测试平台:事故原因涵盖不同损伤维度,事故前的时空线索可调用与自动驾驶、跌倒检测等更广泛安全领域共有的推理能力。我们推出SPRINT(Sports Proactive Risk INference Testbed,运动主动风险推理测试平台),该基准包含2888个真实世界运动视频(2440个事故视频、448个安全对照视频),涉及14项运动和3种环境场景。事故视频带有早期危险线索、事故时间及分层原因的细粒度标注;安全视频经人工验证无事故,用于诊断提示词引发的误报。在不同提示词和时间窗口下评估最先进MLLM,结果显示危险感知与理解能力存在显著差距:最优模型危险信号发出率超95%,但危险原因识别率低于50%。诊断实验进一步表明,明确的危险查询即使在无危险视频上也会引发严重误报。这些发现说明当前MLLM仅具备表层主动安全能力,缺乏稳定、基于原因的早期预警,凸显动态物理环境中可靠主动安全的必要性。数据和代码将在论文录用后开源。

英文摘要:

Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, leaving proactive physical hazard prediction underexplored. Sports provide a well-suited testbed: accident causes span diverse injury dimensions and pre-accident spatiotemporal cues draw on reasoning capabilities shared with broader safety domains such as autonomous driving and fall detection. We introduce SPRINT (Sports Proactive Risk INference Testbed), a benchmark of 2,888 real-world sports videos (2,440 accident, 448 safe controls) spanning 14 sports and 3 environmental settings. Accident videos feature fine-grained annotations of early hazard cues, accident timing, and hierarchical causes; safe videos are manually verified as accident-free and serve to diagnose prompt-induced false alarms. Evaluating state-of-the-art MLLMs under diverse prompts and temporal windows reveals a sharp gap between hazard sensitivity and understanding: the best model exceeds 95% in signaling hazards yet falls below 50% in identifying their causes. Diagnostic experiments further show that explicit danger queries trigger severe false alarms even on hazard-free videos. These findings indicate that current MLLMs exhibit only superficial proactive safety, lacking stable, cause-grounded early warning, and underscore the need for reliable proactive safety in dynamic physical environments. Data and code will be open-sourced upon acceptance.

补充信息

↑