arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EviSafe:基于证据的视觉语言模型安全性评估

EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models

Xuetong Li, Gaofeng Liu

arXiv 2608.23313首次发表:更新:

发表机构

Department of Automation, Shanghai Jiao Tong University(上海交通大学自动化系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于证据的VLM安全性评估框架EviSafe及对应基准EviSafeBench,通过三探针协议评估11个VLMs,发现其未因正确多模态原因可靠安全,需开展超越拒绝次数的评估。

AI 中文摘要

视觉语言模型(VLM)安全性基准通常仅评估最终响应:模型是否拒绝、发出警告或服从。这种结果层面的视角无法判断模型是否因正确的多模态原因而安全。看似安全的行为可能反映关键词触发的弃权(不执行)、遗漏视觉危害或对良性敏感输入的过度拒绝。我们引入EviSafe,这是一个用于VLM安全性的基于证据的框架,共同评估面向自然用户的行为、对文本和视觉证据的明确依据,以及对安全关键证据的反事实变化的行为敏感性。EviSafeBench将该框架实例化为一个受控基准,包含1181个黄金图像-文本场景和2452个针对性反事实变体,覆盖8个安全领域和8个风险源类型。每个场景包含黄金安全决策、证据注释、安全响应策略和反事实干预。三探针协议使用自然响应、证据报告和反事实响应提示查询模型,然后使用感知证据的评判器对其进行评分。在11个被评估的VLMs中,自然严重性准确率范围为27.6%至52.8%,宽松诊断一致性范围为6.1%至29.3%,不安全到安全的反事实转换成功率范围为30.4%至58.4%。这些差距表明,被评估的VLMs并非因正确的多模态原因而可靠安全,推动了超越拒绝次数的评估。

英文摘要

Vision-language model safety benchmarks typically evaluate only final responses: whether a model refuses, warns, or complies. This outcome-level view cannot tell whether a model is safe for the right multimodal reason. Safelooking behavior may reflect keyword-triggered refusal, missed visual hazards, or over-refusal of benign-sensitive inputs. We introduce EviSafe, an evidence-grounded framework for VLM safety that jointly evaluates natural user-facing behavior, explicit grounding in textual and visual evidence, and behavioral sensitivity to counterfactual changes in safety-critical evidence. EviSafeBench instantiates the framework as a controlled benchmark with 1,181 gold image-text scenarios and 2,452 targeted counterfactual variants across eight safety domains and eight risk-source types. Each scenario includes a gold safety decision, evidence annotations, a safe-response policy, and counterfactual interventions. The three-probe protocol queries models with natural-response, evidencereporting, and counterfactual-response prompts, then scores them using an evidence-aware judge. Across eleven evaluated VLMs, natural severity accuracy ranges from 27.6% to 52.8%, relaxed diagnostic consistency from 6.1% to 29.3%, and unsafe-to-safe counterfactual transition success from 30.4% to 58.4%. These gaps show that the evaluated VLMs are not reliably safe for the right multimodal reason and motivate evaluation beyond refusal counts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑