arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21151cs.AI

V-DEAL:将视频安全去校准诊断为理解-拒绝耦合失败

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

Zhetong Zhang, Honghao Fu, Miao Xu, Yiwei Wang, Yujun Cai

首次发表
浏览论文内容

中文总结 AI 辅助

研究视频大语言模型安全对齐问题,提出V-DEAL三级诊断框架分析漏洞机制,通过测试发现模型识别有害视频内容有一定准确率,但配对有害视频与良性查询时攻击成功率高,视觉理解激活拒绝倾向弱,还引入提示注入干预方法降低攻击成功率。

中文摘要 AI 辅助

随着视频大语言模型越来越多地部署在实际应用中,确保其安全对齐变得至关重要。我们发现,与良性查询配对的有害视频比与明确有害查询配对的相同视频实现更高的攻击成功率。为了解这种漏洞的潜在机制,我们提出了V-DEAL,一个三级诊断框架,联合分析模型行为、理解和内部表示方面的这种失败。通过逐步排除感知失败并量化模型的内部拒绝倾向,V-DEAL为分析观察到的漏洞的潜在机制提供了新的诊断视角。我们在三个公共基准上测试了六个视频大语言模型,发现模型以超过81%的准确率正确识别有害视频内容,但在有害视频与良性查询配对的情况下,平均攻击成功率仍达到48.33%。隐藏状态分析进一步表明,视觉理解激活的拒绝倾向比文本理解弱。此外,我们引入了一种提示注入干预方法,平均将攻击成功率降低48.24个百分点,并实现了与基于微调的先前方法相当的性能,为解决视频大语言模型中的此类安全风险提供了有效且实用的手段。

英文摘要

As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuitively, we find that harmful videos paired with benign queries achieve higher attack success rates than the same videos paired with explicitly harmful queries. To understand the underlying mechanism of this vulnerability, we present V-DEAL, a three-level diagnostic framework that jointly analyzes this failure across model behaviour, understanding, and internal representations. By progressively ruling out perception failure and quantifying the model's internal refusal tendency, V-DEAL provides a new diagnostic perspective for analyzing the underlying mechanism of the observed vulnerability. We tested six Video LLMs on three public benchmarks and observed that models correctly recognize harmful video content with over 81\% accuracy, yet the average attack success rate still reaches 48.33\% under the condition pairing harmful videos with benign queries. Hidden-state analysis further shows that visual understanding activates a weaker refusal tendency than textual understanding. Furthermore, we introduce a prompt injection intervention method that reduces attack success rates by an average of 48.24 percentage points and achieves performance comparable to prior fine-tuning-based methods, providing an effective and practical means to address such safety risks in Video LLMs.

发表机构

  • University of Queensland(昆士兰大学)
  • University of California, Merced(加州大学默塞德分校)

机构由 AI 辅助整理,请以论文原文为准。

↑