arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

输入级防御能否迁移到针对VideoLLMs的观测级攻击?

Do Input-Level Defenses Transfer to Observation-Level Attacks on VideoLLMs?

Bangshuo Zhu, Wei Song, Yuxin Cao, Yuezhong Wu, Zhiquan Liu, Yuekang Li, Jingling Xue

arXiv 2609.08331首次发表:更新:

发表机构

University of New South Wales; Griffith University; National University of Singapore; Fuzhou University; Jinan University(新南威尔士大学; 格里菲斯大学; 新加坡国立大学; 福州大学; 暨南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出DefTEval框架,系统评估输入级防御对视频大模型观测级攻击的迁移效果,发现其保护有限且失效,需系统级鲁棒性机制保障安全。

AI 中文摘要

视频大语言模型(VideoLLMs)正越来越多地部署在内容审核和视频分析等安全关键应用中。为了高效处理长视频,VideoLLMs依赖帧采样、令牌压缩和模态融合,这些共同构成一个观测流水线,将原始视频缩减为紧凑的内部表示。最近的观测级攻击利用该流水线阻止模型感知有害内容,然而目前尚无专门针对此威胁设计的防御措施。我们提出了DefTEval,一个受控评估框架,系统性地评估作用于已采样帧像素内容的输入级对抗防御能否缓解观测级攻击。在五种VideoLLMs、十一种代表性防御和五种攻击类型中,我们发现输入级防御提供的保护有限且不一致,有害内容检测率经常接近零。关键的是,即使攻击在每一采样帧中嵌入有害信号,防御也失效,这表明瓶颈不仅在于采样遗漏,还在于进入模型的信号被抑制。令牌压缩丢弃局部特征,模态融合系统性地降低被弱化的视觉信号的权重。此外,防御有效性主要由模型架构而非防御方法本身决定,且检测率在不同内容类别间差异巨大,暴露了时间推理的结构性弱点。这些发现表明,保障VideoLLMs安全需要系统级鲁棒性机制,涵盖采样感知的覆盖保证、安全相关特征的令牌级保留以及模态平衡的融合。

英文摘要

Video Large Language Models (VideoLLMs) are increasingly deployed in safety-critical applications such as content moderation and video analytics. To process long videos efficiently, VideoLLMs rely on frame sampling, token compression, and modality fusion, which together form an observation pipeline that reduces the raw video to a compact internal representation. Recent observation-level attacks exploit this pipeline to prevent the model from perceiving harmful content, yet no defense has been explicitly designed for this threat. We introduce DefTEval, a controlled evaluation framework that systematically assesses whether input-level adversarial defenses, which operate on the pixel content of already-sampled frames, can mitigate observation-level attacks. Across five VideoLLMs, eleven representative defenses, and five attack types, we find that input-level defenses offer limited and inconsistent protection, with harmful detection rates frequently near zero. Critically, defenses fail even against attacks that embed harmful signals in every sampled frame, indicating that the bottleneck extends beyond sampling omission to the suppression of signals that do enter the model. Token compression discards localized features, and modality fusion systematically down-weights weakened visual signals. Furthermore, defense effectiveness is dominated by model architecture rather than by the defense method itself, and detection rates vary drastically across content categories, exposing structural weaknesses in temporal reasoning. These findings demonstrate that securing VideoLLMs requires system-level robustness mechanisms spanning sampling-aware coverage guarantees, token-level preservation of safety-relevant features, and modality-balanced fusion.

CommentsPreprint. Under review at IEEE Transactions on Dependable and Secure Computing. 13 pages, 1 figure, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑