arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

危险还是异常?评估用于理解危险和差异的视觉语言模型

Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

Murali Indukuri, Mohammad Eskandari, Sree Nitya Kollu, Stephanie Lukin, Cynthia Matuszek

arXiv 2607.18325首次发表:更新:

发表机构

Interactive Robotics and Language Lab, University of Maryland Baltimore County; DEVCOM Army Research Laboratory(马里兰大学巴尔的摩县分校交互式机器人与语言实验室; 陆军研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现代安全关键系统中VLM的评估问题,通过明确区分危险和异常,在两个数据集及多种提示策略下评估多个先进VLM,发现其常误判,明确区分能提供更具信息性评估并揭示失败模式。

AI 中文摘要

现代安全关键系统越来越依赖人机交互来降低灾害风险并在紧急情况下支持决策。视觉语言模型(VLM)在这些场景中很有前景,因为它们可以解释复杂场景并传达安全相关信息,但仍需仔细评估以确保可靠的安全推理。当前评估常将危险识别设为二元决策,不清楚模型是识别真正物理危险还是仅对异常场景元素做出反应。我们通过明确区分危险和异常,并分别识别危险和异常状态来解决此限制。我们在两个数据集和多种提示策略上评估了几个先进的VLM,以测试这种区分是否改变模型行为。结果表明VLM常将异常误解为危险,揭示了对上下文不规则性作为危险代理的过度依赖。我们还表明明确区分异常和危险能对VLM安全推理提供更具信息性的评估,并揭示二元安全判断可能掩盖的失败模式。我们的公共数据集可在Roboflow上获取。

英文摘要

Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies. Vision-Language Models (VLMs) are promising for these settings because they can interpret complex scenes and communicate safety-relevant information, but they still require careful evaluation to ensure reliable safety reasoning. In particular, current evaluations often frame danger recognition as a binary decision (Safe/Unsafe), making it unclear whether a model is identifying true physical hazards or merely reacting to unusual scene elements. We address this limitation by introducing an explicit distinction between hazard and anomaly, and by separately recognizing hazardous and anomalous states. We evaluate several state-of-the-art VLMs across two datasets and multiple prompting strategies to test whether this distinction changes model behavior. Our results show that VLMs frequently misinterpret anomalousness as hazardousness, revealing an over-reliance on contextual irregularity as a proxy for danger. We further show that explicitly separating anomaly from hazard provides a more informative evaluation of VLM safety reasoning and exposes failure modes that binary safety judgments can obscure. Our public dataset is available on Roboflow https://app.roboflow.com/vlm-in-context-anomaly-and-hazard-detection/camera-ready-roman-ds.

Comments8 pages, accepted to RO-MAN 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑