arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种模态遗忘一切:增强视觉语言模型中的跨模态遗忘

One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models

Sudharshan Balaji, Yili Ren, Guangjing Wang, Yimin Chen, Ning Wang

arXiv 2607.16442首次发表:更新:

发表机构

University of South Florida; University of Massachusetts Lowell(南佛罗里达大学; 马萨诸塞大学洛厄尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言模型中跨模态遗忘转移问题,通过对三种架构研究发现其不对称不完整。提出CrossInf策略,聚焦关键transformer块遗忘,减少转移差距,提升鲁棒性,经人工评估验证,还用CKA分析浅层遗忘。

AI 中文摘要

机器遗忘被广泛用于从大语言模型中去除有害知识。然而,现代视觉语言模型(VLM)同时处理文本和视觉输入,引发了一个基本安全问题:一种模态的遗忘是否能转移到另一种模态?我们首次对三种VLM架构(LLaVA - 1.5(MLP投影)、InstructBLIP(Q - Former)和IDEFICS(门控交叉注意力))进行了系统的双向跨模态遗忘转移研究。发现遗忘能跨模态转移,但不对称且不完整。文本遗忘在某些情况下能强烈转移到视觉,但在排版攻击下不稳健,先前遗忘的知识易恢复。为此提出CrossInf,一种影响引导的缓解策略。它聚焦于对跨模态泛化影响最大的transformer块进行遗忘,在强融合架构中减少了一半以上的转移差距,提高了排版攻击下的鲁棒性,攻击成功率降至近零。还通过人工评估验证了结果,并用CKA分析了浅层遗忘。

英文摘要

Machine unlearning is widely used to remove hazardous knowledge from large language models. Modern Vision-Language Models (VLMs), however, process both text and visual inputs, raising a fundamental security question: does unlearning in one modality transfer to the other? We present the first systematic, bidirectional study of cross-modal unlearning transfer across three VLM architectures: LLaVA-1.5 (MLP projection), InstructBLIP (Q-Former), and IDEFICS (gated cross-attention). We find that unlearning transfers across modalities, but the transfer is asymmetric and incomplete. In some cases, text unlearning strongly transfers to vision. However, this robustness is not preserved under typographic attacks that manipulate the visual presentation of text. Under such attacks, previously unlearned knowledge can be readily recovered, indicating shallow unlearning. To address the transfer gap and shallow robustness, we propose \textsc{CrossInf}, an influence-guided mitigation strategy. Motivated by the observation that different model components contribute unequally to cross-modal transfer, \textsc{CrossInf} focuses unlearning on transformer blocks that most influence cross-modal generalization. It reduces the transfer gap by more than half in architectures with strong fusion, while preserving model utility. It also improves robustness under typographic attacks, reducing the attack success rate to near zero. We further conduct human evaluation with three annotators ($κ{=}0.77$) to validate our findings. Finally, we analyze shallow unlearning using Centered Kernel Alignment (CKA), providing insights into the observed transfer behavior and robustness limitations.

Comments18 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑