视觉令牌压缩增强多模态大语言模型的鲁棒性
Visual Token Compression Enhances Robustness of MLLMs
- Hefei University of Technology(合肥工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究如何增强多模态大语言模型的鲁棒性,核心方法是通过测量视觉令牌与语言特征空间的距离来识别并剪枝分布外视觉令牌,该方法能有效防御越狱攻击、减轻幻觉,提升模型在通用数据集上的表现。
AI中文摘要:
在本文中,我们首次表明视觉令牌剪枝可增强多模态大语言模型(MLLM)的鲁棒性,减轻越狱攻击和幻觉等漏洞。鉴于视觉和语言模态无法完美对齐,未对齐的视觉令牌可能充当分布外(OOD)输入,导致不可预测的输出并引入潜在漏洞。基于此,我们旨在通过在鲁棒剪枝层减少OOD视觉令牌来增强模型对越狱和幻觉的鲁棒性,同时还能降低推理成本。具体而言,我们测量每个视觉令牌与语言特征空间之间的距离,将距离大的视觉令牌识别为OOD令牌并进行迭代剪枝。为证明该方法的有效性,我们在七个不同的流行基准上进行评估。值得注意的是,我们的方法在防御越狱攻击方面平均提高了13.29%,在减轻幻觉方面始终具有有竞争力的性能,并在MME等通用数据集上保持良好结果。
英文摘要:
In this paper, we show for the first time that visual token pruning enhances the robustness of Multimodal Large Language Models (MLLMs), mitigating vulnerabilities such as jailbreak attacks and hallucinations. Given that vision and language modalities cannot be perfectly aligned, the misaligned visual tokens might act as out-of-distribution (OOD) inputs, leading to unpredictable outputs and introducing potential vulnerabilities. Building on this insight, we aim to enhance model robustness against jailbreaks and hallucinations by reducing OOD visual tokens at robust-pruning layers, while also reducing inference cost as a side benefit. Specifically, we measure the distance between each visual token and the language feature space. Then, visual tokens with large distances are identified as OOD tokens, which can be iteratively pruned. To demonstrate the effectiveness of our method, we evaluate it on seven diverse popular benchmarks. Notably, our method yields an average improvement of 13.29\% in defending jailbreak attacks, consistently achieves competitive performance in mitigating hallucinations, and maintains strong results on general datasets like MME.