arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

仍在,却已不见:揭示大型视觉语言模型中的压缩诱导风险

Still There, No Longer Seen: Exposing Compression-Induced Risk in Large Vision-Language Models

Qiankun Li, Yuechen Zhang, Bowen Chen, Shilinlu Yan, Zhenhong Zhou, Kun Wang, Li Sun

arXiv 2609.35002首次发表:更新:

发表机构

Nanyang Technological University; Beijing University of Posts and Telecommunications(南洋理工大学; 北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大型视觉语言模型,提出压缩诱导风险攻击CIRA,通过操纵令牌优先级在受限访问下实现高压缩特定失败率,并验证了配对评估的必要性。

AI 中文摘要

视觉令牌压缩降低了大型视觉语言模型(LVLMs)的推理成本。然而,整体的鲁棒性度量无法揭示某个特定的对抗性失败是由压缩引起的,还是继承自底层模型。我们将压缩特定失败(CSF)定义为在完整令牌推理下保持正确,但在压缩后失败的对抗性输入,从而将压缩诱导风险视为一个配对失败归因问题。在一个受控的诊断队列中,反事实表明保留集的分配因果性地改变了压缩后的正确性,并揭示了恢复与移位证据中的表示漂移之间的负相关关系。受这些发现的启发,我们提出了CIRA,一种针对大型视觉语言模型的压缩诱导风险攻击。在视觉编码器白盒设置下,CIRA通过编码器侧目标优化图像扰动,这些目标在候选压缩预算之间操纵令牌优先级,同时保留移位证据。CIRA不使用下游问题或标签,也不需要访问语言模型、部署的压缩器或确切的压缩预算。在四个预算下评估的12个数据集-压缩器设置中,CIRA实现了平均CSFR为20.35%,同时将完整令牌攻击成功率限制在6.92%,在额外的LVLM系列上也表现出类似行为。一种跨视图选择稳定防御显著抑制了CIRA,尽管自适应CIRA部分恢复了其有效性。这些结果表明,压缩特定失败在受限访问下仍然存在,并支持对完整令牌和压缩推理进行配对评估,以将风险归因于视觉令牌压缩。

英文摘要

Visual token compression reduces the inference cost of Large Vision-Language Models (LVLMs). However, aggregate robustness measures do not reveal whether a particular adversarial failure is induced by compression or inherited from the underlying model. We define a compression-specific failure (CSF) as an adversarial input that remains correct under full-token inference but fails after compression, casting compression-induced risk as a paired failure attribution problem. Within a controlled diagnostic cohort, counterfactuals show that retained-set allocation causally changes compressed correctness and reveal a negative association between recovery and representation drift in displaced evidence. Motivated by these findings, we propose CIRA, a Compression-Induced Risk Attack for Large Vision-Language Models. Under a vision-encoder white-box setting, CIRA optimizes image perturbations through encoder-side objectives that manipulate token priorities across candidate compression budgets while preserving displaced evidence. CIRA uses no downstream questions or labels and requires no access to the language model, deployed compressor, or exact compression budget. Across 12 dataset-compressor settings evaluated at four budgets, CIRA achieves a mean CSFR of 20.35% while limiting full-token attack success to 6.92%, with similar behavior on additional LVLM families. A cross-view selection-stabilization defense substantially suppresses CIRA, although Adaptive CIRA partially restores its effectiveness. These results show that compression-specific failures persist under restricted access and support paired evaluation of full-token and compressed inference for attributing risk to visual-token compression.

Comments29 pages, 11 figures, 13 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑