并非所有补丁都同样可遗忘:视觉-语言模型中的空间局部化领域遗忘
Not All Patches Are Equally Forgettable: Spatially Localized Domain Unlearning in Vision-Language Models
- Indian Institute of Science Education and Research Bhopal(印度科学教育与研究学院博帕尔分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出两阶段补丁选择框架,通过定位并抑制对遗忘领域预测贡献大的补丁,实现视觉-语言模型的空间局部化领域遗忘,在多个基准上改善遗忘-保留权衡并提升保留领域识别性能。
AI中文摘要:
预训练的视觉-语言模型(VLM)即使没有额外训练,也展现出强大的跨领域识别性能。然而,这种鲁棒性也可能保留不良的领域特定行为,因为领域相关信息和语义信息常常在学习到的表示空间内纠缠在一起,使得选择性领域遗忘具有挑战性。现有方法通常通过潜在空间解缠和提示或特征级干预来解决这一问题,而不直接归因和减弱单个补丁令牌的贡献。然而,我们在此提出,与其均匀抑制整个表示,利用视觉变换器的空间结构来定位并抑制对遗忘领域预测贡献不成比例的补丁区域可能更为有效。对遗忘领域预测有强烈影响的补丁可能对语义识别并非同等重要,这表明遗忘应根据不同视觉区域的领域贡献来引导。具体而言,我们提出一个两阶段补丁选择框架,首先估计补丁级领域敏感性,然后选择性地减弱那些对遗忘领域预测的贡献强于其语义效用的补丁。我们在Office-Home、Mini DomainNet和DomainNet上评估了我们的框架。实验结果表明,与先前方法相比,遗忘-保留权衡得到改善,同时保留领域识别性能提升高达3.8%。在视觉重叠和未见领域设置下的额外评估进一步证明了在分布偏移下的鲁棒性提升。
英文摘要:
Pre-trained vision-language models (VLMs) exhibit strong cross-domain recognition performance even without additional training. However, this robustness can also preserve undesirable domain-specific behavior, as domain-related and semantic information often remain entangled within the learned representation space, making selective domain unlearning challenging. Existing approaches typically address this problem through latent-space disentanglement and prompt- or feature-level interventions, without directly attributing and attenuating individual patch-token contributions. However, here we suggest that rather than uniformly suppressing the full representation, it may be more effective to exploit the spatial structure of vision transformers to localize and suppress patch regions that contribute disproportionately to forget-domain prediction. Patches that strongly influence forget-domain prediction may not be equally important for semantic recognition, suggesting that forgetting should be guided according to the domain contribution of different visual regions. Specifically, we propose a two-stage patch-selective framework that first estimates patch-level domain sensitivity and then selectively attenuates patches whose contribution to forget-domain prediction is stronger than their semantic utility. We evaluate our framework on Office-Home, Mini DomainNet, and DomainNet. Experimental results demonstrate improved forgetting-retention tradeoffs compared to prior methods while improving retained-domain recognition by up to 3.8\%. Additional evaluations under visually overlapping and unseen-domain settings further demonstrate improved robustness under distribution shift.