发表机构
AIM Intelligence; Seoul National University(AIM智能研究所; 首尔国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出EgoSafetyBench,一个包含1200个机器人视角场景的自我中心视频基准,用于评估视觉语言模型在流式安全监控中的能力,通过情景和视觉通道两个轨道测试模型识别危险和应对误导性文本的鲁棒性。
AI 中文摘要
视觉语言模型(VLM)现被提议作为家庭和工厂中具身代理的运行时安全卫士。一个可部署的卫士必须捕捉真正不安全的情况,同时避免对常规但表面令人警觉的活动进行不必要的干预,这一区别被二元安全基准所模糊。我们引入了EgoSafetyBench,一个包含1200个机器人视角场景的自我中心视频基准,以半秒粒度注释,用于评估VLM作为流式卫士在两个轨道上的表现。情景轨道(800个场景)涵盖四个类别,从常规和安全但可疑的场景到明显和上下文相关的危险。视觉通道轨道(400个场景)针对场景内文本——场景中可见的标志、贴纸或标签——这些文本可能歪曲物理情况,将每个误导性标志与真实版本配对,以测试卫士是否将文本标记为误导性以及文本是否损害其物理安全判断。两个轨道都使用对比阶梯:几乎相同的场景仅在一个可见的决定性线索上不同,因此正确的判断必须依赖于该线索而非整体场景类型。我们评估了十个开源和闭源VLM。我们发现,虽然卫士能可靠地识别包含危险的视频,但它们经常错过特定的危险时刻,特别是上下文相关的危险。此外,误导性的场景内文本降低了所有测试的卫士的性能:脆弱的模型错过了多达三分之一的危险,而鲁棒的模型对安全内容过度干预。匹配对照表明,表面上的安全鲁棒性往往反映的是不加区分的警报而非真正的物理推理。
英文摘要
Vision-language models (VLMs) are increasingly proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genuinely unsafe situations while avoiding unnecessary intervention on routine but superficially alarming activity, a distinction obscured by binary safety benchmarks. We introduce EgoSafetyBench, an egocentric video benchmark of 1,200 robot-view scenarios annotated at half-second granularity, with two evaluation tracks. The situational track (800 scenarios) spans routine, safe-but-suspicious, obvious-hazard, and contextual-hazard scenes. The visual-channel track (400 scenarios) tests whether misleading in-scene text corrupts physical-safety judgments, using matched truthful controls. Both tracks use contrastive ladders: near-identical scenarios differing in a single visible deciding cue, forcing predictions to hinge on that cue. Across ten open- and closed-source VLMs, we find that guards often recognize videos containing hazards yet miss the specific hazardous moments, especially for contextual hazards. Misleading in-scene signs further degrade all tested guards: vulnerable models miss up to a third of hazards, while seemingly robust models often over-intervene on safe content. Matched controls show that apparent robustness can reflect indiscriminate alarming rather than true physical reasoning. A 20-clip real-video sanity check further shows that model rankings transfer beyond synthetic rendering, with Spearman \r{ho} = 0.87 for hazard miss rate and 0.94 for visual-channel mismatch recall.