AI 中文总结
本文针对跨模态知识遗忘迁移研究缺口,推出UNLINK-VL基准,实验发现跨模态遗忘存在不对称性,纯文本遗忘迁移效果差,凸显跨模态遗忘与评估的必要性。
AI 中文摘要
视觉-语言模型(Vision-Language Models, VLMs)与大语言模型(Large Language Models, LLMs)类似,可能会在预训练语料中记忆敏感、受版权保护或有害的知识,移除此类知识对于构建可信AI系统至关重要。然而现有研究主要聚焦于单模态内的遗忘,尽管近期研究已开始探索遗忘过程中的跨模态一致性,但真实世界知识遗忘的跨模态迁移仍未得到充分研究。为填补这一空白,我们推出UNLINK-VL——一个面向VLMs跨模态知识遗忘的真实世界基准。在无法获取原始遗忘语料与保留语料的事后遗忘场景下,UNLINK-VL选取视觉可识别的真实世界实体作为遗忘目标,将其与对应图像及源自Wikidata的一跳、多跳事实关联。该基准包含四个互补子集,分别评估目标知识的直接遗忘、遗忘通过关联知识的传播、相关非目标知识的保留以及对语义等价查询的鲁棒性。我们在纯文本与多模态遗忘场景下训练模型,并在文本、视觉及跨模态场景中评估遗忘效果与保留效用。大量实验揭示了跨模态迁移的显著不对称性:多模态遗忘在文本评估下仍保持有效,而纯文本遗忘向视觉及跨模态场景的迁移效果不佳;同时,所评估的方法在很大程度上保留了模型的通用能力。这些发现表明,仅依赖单模态内评估(尤其是纯文本评估)可能会大幅高估VLMs中知识遗忘的效果,凸显了跨模态遗忘与评估的必要性。
英文摘要
Vision-Language Models (VLMs), like Large Language Models (LLMs), may memorize sensitive, copyrighted, or harmful knowledge from their pretraining corpora. Removing such knowledge is essential for building trustworthy AI systems. However, existing studies primarily focus on forgetting within individual modalities. Although recent work has begun to explore cross-modal consistency in unlearning, the cross-modal transfer of real-world knowledge unlearning remains insufficiently studied. To address this gap, we introduce UNLINK-VL, a real-world benchmark for cross-modal knowledge unlearning in VLMs. Under a post-hoc unlearning setting in which the original forget and retain corpora are unavailable, UNLINK-VL selects visually identifiable real-world entities as unlearning targets and associates them with corresponding images and one-hop and multi-hop facts derived from Wikidata. The benchmark comprises four complementary subsets that evaluate direct forgetting of target knowledge, the propagation of forgetting through relational knowledge, the preservation of related non-target knowledge, and robustness to semantically equivalent queries. We train models under text-only and multimodal unlearning settings and evaluate forgetting effectiveness and retained utility across textual, visual, and cross-modal scenarios. Extensive experiments reveal a pronounced asymmetry in cross-modal transfer: multimodal unlearning remains effective under textual evaluation, whereas text-only unlearning transfers poorly to visual and cross-modal scenarios. Meanwhile, the evaluated methods largely preserve the models' general capabilities. These findings demonstrate that relying solely on intra-modal evaluation, particularly text-only evaluation, may substantially overestimate the effectiveness of knowledge unlearning in VLMs, underscoring the need for cross-modal unlearning and evaluation.