发表机构
Guangzhou University; University of Macau(广州大学; 澳门大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出联合像素-提示优化(JPPO)框架,将提示与图像联合优化,对VLM实现显著成本放大,揭示多模态服务防御的结构性盲点。
AI 中文摘要
针对自回归视觉语言模型(VLM)的资源耗尽攻击通常假设单模态威胁模型,将图像分支视为主要优化面,同时保持用户可见的提示固定不变。即使是最近的循环中心变体也局限于这种单通道范式,未将可用性利用作为跨模态优化问题在联合可控输入面上进行探索。我们提出了联合像素-提示优化(JPPO),这是首个将可见提示提升为与图像扰动同等重要的对抗变量的复合对抗框架。在受限的联合输入威胁模型下,JPPO在像素和提示两个面上执行耦合的分阶段优化。这产生了协同的成本放大效应,在机制上区别于循环依赖的故障,在我们的实验中循环发生率可忽略不计。在8/255无穷范数预算下,对MS COCO和ImageNet上的五个开源VLM系列进行评估,JPPO在Qwen2.5-VL-7B上实现了超过4.6倍的延迟和5.3倍的能量放大,在BLIP-2上实现了超过36.6倍的延迟和32.7倍的能量放大。在直接比较的基线中,这代表了最强的成本放大,同时所需的优化迭代次数大幅减少。消融实验证实,这种放大源于多模态协调,而非提示长度或孤立模态。这些发现揭示了当前VLM服务防御中的结构性盲点,促使将成本感知的鲁棒性评估作为多模态部署的一等安全要求。
英文摘要
Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal threat models, treating the image branch as the primary optimization surface while holding user-visible prompts fixed. Even recent loop-centric variants remain confined to this single-channel paradigm, leaving the exploitation of availability unexplored as a cross-modal optimization problem over jointly controllable input surfaces. We introduce Joint Pixel-Prompt Optimization (JPPO), the first compound adversarial framework elevating the visible prompt to a first-class adversarial variable alongside image perturbations. Under a restricted joint-input threat model, JPPO performs coupled, stagewise optimization over both the pixel and prompt surfaces. This produces synergistic cost amplification, mechanistically distinct from loop-dependent failures, exhibiting negligible loop incidence in our experiments. Evaluating five open-source VLM families on MS COCO and ImageNet under an 8/255 infinity-norm budget, JPPO achieves over 4.6x latency and 5.3x energy amplification on Qwen2.5-VL-7B, and over 36.6x latency with 32.7x energy amplification on BLIP-2. This represents the strongest cost amplification among directly compared baselines while requiring substantially fewer optimization iterations. Ablations confirm this amplification arises from multimodal coordination rather than prompt length or isolated modalities. These findings reveal structural blind spots in current VLM serving defenses, motivating cost-aware robustness evaluation as a first-class security requirement for multimodal deployments.
CommentsExtended version of the paper accepted by ACM CCS 2026