arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向细粒度遗忘:多模态大语言模型的属性遗忘

Toward Fine-Grained Forgetting:Attribute Unlearning for Multimodal Large Language Models

Junkai Lin, Junkai Chen, Siqi Hou, Yuhao He, Ruiqi Liu, Chenhan Jin, Shengze Xu, Tieyong Zeng

arXiv 2608.01008首次发表:更新:

AI 中文总结

针对多模态大语言模型(MLLMs)属性遗忘的细粒度需求,提出轻量级训练无关框架CLRP,通过激活修补定位因果层并应用保留感知投影,实现目标属性遗忘同时保留同一身份信息,在多种MLLMs上验证了有效性。

AI 中文摘要

多模态大语言模型(Multimodal Large Language Models,MLLMs)具备强大的视觉-语言能力,但也可能记忆并泄露敏感信息。机器遗忘旨在从模型中移除指定知识,且无需从头重新训练,同时保留模型的通用实用性。现有的隐私导向基准主要采用个人资料级别的删除,而实际需求往往更细粒度:模型应遗忘指定属性,同时保留同一身份的非敏感信息。因此,我们将多模态大语言模型的属性遗忘作为一项细粒度任务,并构建了涵盖长文本、数值和短文本目标、多种遗忘率以及多样问题类型的基准。我们的评估显示,目标属性与保留属性共享身份特定信息和视觉证据,这使得选择性遗忘易受残留泄露或附带性能下降的影响;相应地,现有方法在该场景下的遗忘-保留权衡表现不稳定。为应对这一挑战,我们提出了因果定位与保留感知投影(Causal Localization and Retain-Aware Projection,CLRP),这是一个轻量级、无需训练的框架。CLRP通过激活修补识别出因果介导目标属性泄露的层,随后应用保留感知投影,在移除目标属性子空间的同时保留同一身份的证据。在多个广泛使用、具有不同架构和参数规模的MLLMs上进行的实验证明了CLRP的有效性。

英文摘要

Multimodal large language models (MLLMs) exhibit strong vision--language capabilities but may also memorize and disclose sensitive information. Machine unlearning seeks to remove designated knowledge without retraining from scratch while preserving general utility. Existing privacy-oriented benchmarks primarily adopt profile-level deletion, whereas practical requests are often finer grained: a model should forget a specified attribute while retaining non-sensitive information about the same identity. We therefore introduce attribute-level MLLM unlearning as a finer-grained task and construct a benchmark spanning long-text, numeric, and short-text targets, multiple forget ratios, and diverse question types. Our evaluation reveals that target and retained attributes share identity-specific and visual evidence, making selective forgetting susceptible to residual leakage or collateral degradation; accordingly, existing methods exhibit unstable forgetting--retention trade-offs in this setting. To address this challenge, we propose Causal Localization and Retain-Aware Projection (CLRP), a lightweight training-free framework. CLRP uses activation patching to identify the layer that causally mediates target-attribute disclosure, then applies a retain-aware projection that removes the target-attribute subspace while preserving same-identity evidence. Experiments across multiple widely used MLLMs with distinct architectures and parameter scales demonstrate the effectiveness of CLRP.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑