多模态安全评估应衡量超越分类的可控性
Multimodal Safety Evaluation Should Measure Controllability Beyond Classification
浏览论文内容
中文总结 AI 辅助
本文提出多模态安全评估应报告可控性画像,超越传统分类,通过隐性毒性测试证明内部读出与可控性可能分歧,建议基准增加干预测试。
中文摘要 AI 辅助
VLM(视觉语言模型)的安全性通常通过输入级和输出级分类来评估。这种分类是必要的,但它并不能揭示安全状态在模型内部是否可访问或可控。我们认为,多模态安全评估因此应在行为分类之外报告一个“可控性画像”,将表征级可检测性、跨模态特异性、干预敏感性和良性保持选择性区分开来。以隐性毒性作为压力测试案例,我们使用稀疏特征分解在LlavaGuard和Qwen3.5上实例化该画像。LlavaGuard提供了局部化操作手柄,但其良性保持干预范围狭窄,且下游安全性提升有限;而Qwen3.5支持强大的表征级读出,但在所测试的操作算子下缺乏可比的选择性控制机制。这些结果表明,内部读出与可控性可能产生分歧。未来的多模态安全基准因此不仅应报告行为安全指标,还应报告与安全相关的内部信号是否能在经过验证的操作范围内进行干预测试和控制。
英文摘要
VLM safety is commonly evaluated through input- and output-level classification. Such classification is necessary, but it does not reveal whether a safety state is accessible or controllable inside the model. We argue that multimodal safety evaluation should therefore report a \emph{controllability profile} alongside behavioral classification, separating representation-level detectability, cross-modal specificity, intervention sensitivity, and benign-preserving selectivity. Using implicit toxicity as a stress case, we instantiate this profile on LlavaGuard and Qwen3.5 with sparse feature decompositions. LlavaGuard admits localized handles with a narrow benign-preserving intervention range and modest downstream safety gains, whereas Qwen3.5 supports strong representation-level readout but no comparable selective-control regime under the tested operators. These results show that internal readout and controllability can diverge. Future multimodal safety benchmarks should therefore report not only behavioral safety metrics, but also whether safety-relevant internal signals can be intervention-tested and controlled within a validated operating range.
发表机构
- KAIST(韩国科学技术院)
- AIM Intelligence
- Seoul National University(首尔大学)
- KT Corporation(KT公司)
- University of British Columbia(不列颠哥伦比亚大学)
- Vector Institute(向量研究所)
机构由 AI 辅助整理,请以论文原文为准。