arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

理解并缓解视觉语言模型中由Token剪枝引发的安全漏洞

Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs

Shuailong Wang, Xinyu Lyu, Shengming Yuan, Jingkuan Song, Heng Tao Shen, Lianli Gao

arXiv 2610.09703首次发表:更新:

发表机构

University of Electronic Science and Technology of China; Southwestern University of Finance and Economics; Tongji University; Shanghai Innovation Institute(电子科技大学; 西南财经大学; 同济大学; 上海创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文首次系统评估Token剪枝对视觉语言模型安全性的影响,发现其可引发恶意放大漏洞,并提出即插即用的安全感知剪枝(SAP)机制,通过识别恶意锚点、恢复良性Token和重分配注意力,在不牺牲效率的前提下将攻击成功率降低高达62%。

AI 中文摘要

Token剪枝通过移除冗余的视觉Token来加速视觉语言模型,但其安全性影响仍未得到充分探索。在本工作中,我们首次对Token剪枝机制进行了全面的安全性评估,并发现:大多数剪枝策略随着剪枝比例的增加会显著降低安全性,而基于查询的压缩则表现出相反的趋势,极端剪枝(高达99.8%)反而意外地提升了模型安全性。这种鲜明对比引发了一个关键问题:不同的Token剪枝策略如何重塑模型的安全行为,以及是否有可能在不牺牲加速效果的前提下提升安全性?为了回答这一问题,我们识别出一种未被认知的机制,称为“剪枝引发的恶意放大”,即移除背景Token会产生副作用:迫使模型的注意力集中到前景中少数保留的恶意锚点上,从而在越狱攻击下无意中放大其毒性语义。为解决这一问题,我们提出了一种推理时、即插即用的安全感知剪枝(SAP)机制,通过三个步骤来抵消这种主导效应:(1)识别恶意锚点,(2)恢复被剪枝的良性Token,以及(3)将过度集中的注意力从恶意锚点重新分配到良性Token上。在三个安全基准和四个效用基准上的大量实验表明,SAP能够缓解剪枝引发的安全漏洞,即在不影响效率或效用的情况下,将攻击成功率(ASR)降低高达62%。

英文摘要

Token-Pruning accelerates Vision-Language Models by removing redundant visual tokens, yet its safety implications remain underexplored. In this work, we present the first comprehensive safety evaluation of Token-Pruning mechanisms and find that: most pruning strategies significantly degrade safety as pruning ratios increase, whereas Query-based Compression shows the opposite, with extreme pruning (up to 99.8%), unexpectedly improves model safety. This sharp contrast prompts a key question: How do different Token-Pruning strategies reshape model safety behavior, and is it possible to enhance safety without sacrificing acceleration? To answer this, we identify an unrecognized mechanism, termed Pruning-Induced Malicious Amplification, where removal of background tokens triggers a side effect: forcing the model's attention to collapse onto a few retained malicious anchors within the foreground, inadvertently amplifying their toxic semantics under jailbreak. To address that, we propose an inference-time and plug-and-play Safety-Aware Pruning (SAP) mechanism that counteracts such dominance via three steps: (1) identifying malicious anchors, (2) restoring pruned benign tokens, and (3) reallocating excessive attention from malicious anchors to benign tokens. Extensive experiments across three safety and four utility benchmarks demonstrate that SAP mitigates pruning-induced vulnerabilities, i.e., reducing ASR by up to 62%, without compromising efficiency or utility.

CommentsAccepted at ICML 2026. 16 pages

Journal refProceedings of the 43rd International Conference on Machine Learning, PMLR 306:128879-128894, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑