arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01962cs.LGcs.CV

SIEVE:面向视觉-语言模型遗忘的选择性注意力值抑制

SIEVE: Selective attention-value Suppression for Vision-Language Models Unlearning

Si Qi Goh, Cap Dang Xuan Kiet, Tat-Jen Cham, Kwok-Yan Lam

首次发表
浏览论文内容

中文总结 AI 辅助

SIEVE通过抑制遗忘样本的注意力值并匹配保留样本表示,实现了视觉-语言模型中对个人可识别信息的选择性遗忘,在保持保留知识的同时达到最先进性能。

中文摘要 AI 辅助

视觉-语言模型(VLM)将视觉身份与传记信息关联的能力,产生了在保留关于同一人的允许知识的同时,选择性遗忘个人可识别信息(PII)的需求。这一设置具有挑战性,因为敏感信息和保留信息可能共享相同的视觉输入和中间表示。我们提出了SIEVE,一个简单而有效的选择性VLM遗忘框架。SIEVE直接正则化注意力值表示,同时控制模型输出。SIEVE将遗忘样本的注意力值抑制为常数零,同时通过将保留样本的表示与冻结的参考模型匹配来保持其表示。这些目标与序列级别的遗忘和保留监督相结合,使得在不大幅影响保留知识的情况下实现有针对性的遗忘。大量实验表明,SIEVE在多种模型模态设置下的遗忘任务中达到了最先进的性能,同时保持了具有竞争力的保留效用。消融研究进一步表明,值抑制和负交叉熵贡献了互补的遗忘信号,而基于参考的值匹配显著减少了效用退化。这些结果表明,当敏感知识和保留知识密切相关时,注意力值为选择性多模态遗忘提供了一个有效的干预点。

英文摘要

The ability of vision-language models (VLMs) to associate visual identities with biographical information creates a need for selective unlearning of personally identifiable information (PII) while preserving permitted knowledge about the same individual. This setting is challenging because both sensitive and retained information can share the same visual inputs and intermediate representations. We introduce SIEVE, a simple and effective framework for selective VLM unlearning. SIEVE directly regularizes attention-value representations while also controlling model outputs. SIEVE suppresses attention values for forget examples toward a constant zero, while preserving retain-example representations by matching them to a frozen reference model. These objectives are combined with sequence-level forget and retain supervision, enabling targeted forgetting without largely affecting retained knowledge. Extensive experiments show that SIEVE achieves state-of-the-art performance on unlearning with multiple model-modality settings, while maintaining competitive retained utility. Ablation studies further show that value suppression and negative cross-entropy contribute complementary forgetting signals, while reference-based value matching substantially reduces utility degradation. These results demonstrate that attention values provide an effective intervention point for selective multimodal unlearning when sensitive and retained knowledge are closely related.

发表机构

  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑