AI 中文总结
针对未学习多模态大语言模型存在的知识缺口问题,提出带锚定正则化的选择性保护方法,在保障未学习安全性的同时大幅恢复模型响应质量。
AI 中文摘要
机器未学习为移除多模态大语言模型(Multimodal Large Language Models,MLLMs)中的不安全内容提供了有前景的方法,但确保未学习的精准性仍是持续存在的挑战。原因之一是当前MLLM未学习评估范式存在关键盲区:它们通过与遗忘集表征相距较远的基准评估模型效用,无法捕捉知识缺口——即对良性相邻输入的严重性能退化。为探究未学习MLLMs中的知识缺口,我们构建了一个基准,用于捕捉与遗忘集共享通用模式的良性输入上的意外退化,并通过控制实验证实,这些退化是常用方法的系统性后果。此外,为弥合这一差距,我们提出了带锚定正则化的选择性保护(Selective Protection with Anchored Regularization,SPAR),该方法通过锚定激活过滤保护通用模式,同时通过实体抽象增强强化这些模式。我们在SafeEraser上的实验表明,与标准基准低于50%的表现相比,SPAR恢复了超过98%的原始模型响应质量,同时实现了0.00%的攻击成功率和具有竞争力的模型效用。这些结果强调了对可信MLLM未学习进行更细粒度评估的必要性。
英文摘要
Machine unlearning offers a promising approach to remove unsafe content from Multimodal Large Language Models (MLLMs), yet ensuring the precision of unlearning remains a persistent challenge. One reason is that current MLLM unlearning evaluation paradigms suffer from a critical blind spot: they assess model utility through benchmarks whose representations are distant from the forget set, failing to capture knowledge holes---severe degradation on benign adjacent inputs. To probe knowledge holes in unlearned MLLMs, we construct a benchmark that captures unintended degradation on benign inputs sharing generic patterns with the forget set, and confirm through controlled experiments that they are a systematic consequence of commonly used approaches. Furthermore, to bridge this gap, we propose Selective Protection with Anchored Regularization, which protects generic patterns via anchored activation filtering while reinforcing them through entity-abstracted enhancement. Our experiments on SafeEraser demonstrate that SPAR recovers over 98% of vanilla response quality compared to below 50% for standard baselines---while achieving 0.00% attack success rate and competitive model utility. These results underscore the necessity of more fine-grained evaluation for trustworthy MLLM unlearning.