Open-MMUnlearning:统一多模态大语言模型机器遗忘的方法与评估
Open-MMUnlearning: Unifying Methods and Evaluation for MLLM Unlearning
浏览论文内容
中文总结 AI 辅助
本文提出Open-MMUnlearning开源框架,统一多模态大模型遗忘的方法与评估,覆盖多基准、模型和方法,并引入度量元评估,发现GD与MIP-Editor表现最优,BLEU度量最可靠。
中文摘要 AI 辅助
随着多模态大语言模型(MLLM)能力不断增强并被广泛部署,隐私和安全问题日益紧迫。机器遗忘提供了一种解决这些问题的途径,即从训练好的模型中移除指定信息,同时保留不相关的功能。然而,实现和评估协议的碎片化、鲁棒性测试的不完整以及对度量可靠性的有限理解,使得难以系统评估MLLM遗忘的进展。我们推出了Open-MMUnlearning,一个开源、可扩展的框架,通过共享接口和结构化配置整合了目标模型准备、多模态数据处理、遗忘和评估。该框架支持涵盖隐私、安全和版权的五个基准,来自四个模型家族的八个MLLM,以及十二种遗忘方法。其评估套件联合评估遗忘效果、保留效用以及对模型干预、对抗性输入和成员推理攻击的鲁棒性。使用统一的评估协议,我们比较了十种代表性遗忘方法。在此比较中,GD和MIP-Editor并列获得最高总分:GD达到最高的遗忘质量,而MIP-Editor保留了更多的模型效用。我们进一步引入了一种度量元评估协议,该协议使用对目标知识具有受控暴露的模型测试忠实性,并在量化和重新学习下测试鲁棒性。在评估的十三个度量中,BLEU达到最高的综合可靠性得分。KS-Test达到最高的忠实性AUC,但在鲁棒性上表现较差。该框架和这些发现共同支持MLLM遗忘方法的可重复比较和评估可靠性的系统评估。
英文摘要
As multimodal large language models (MLLMs) become more capable and widely deployed, concerns about privacy and safety have become increasingly pressing. Machine unlearning offers one approach to addressing these concerns by removing designated information from trained models while preserving unrelated capabilities. However, fragmented implementations and evaluation protocols, incomplete robustness testing, and limited understanding of metric reliability make progress in MLLM unlearning difficult to assess systematically. We introduce Open-MMUnlearning, an open-source, extensible framework that integrates target-model preparation, multimodal data processing, unlearning, and evaluation through shared interfaces and structured configurations. The framework supports five benchmarks spanning privacy, safety, and copyright, eight MLLMs from four model families, and twelve unlearning methods. Its evaluation suite jointly assesses forgetting effectiveness, retained utility, and robustness to model interventions, adversarial inputs, and membership inference attacks. Using a common evaluation protocol, we compare ten representative unlearning methods. In this comparison, GD and MIP-Editor tie for the highest overall score: GD achieves the highest Forget Quality, while MIP-Editor preserves more Model Utility. We further introduce a metric meta-evaluation protocol that tests faithfulness using models with controlled exposure to target knowledge and robustness under quantization and relearning. Among the thirteen evaluated metrics, BLEU achieves the highest aggregate reliability score. KS-Test attains the highest faithfulness AUC but performs less well on robustness. Together, the framework and these findings support reproducible comparison of MLLM unlearning methods and systematic assessment of evaluation reliability.
发表机构
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- University of the Chinese Academy of Sciences(中国科学院大学)
- ByteDance(字节跳动)
- The Chinese University of Hong Kong(香港中文大学)
- University of Cambridge(剑桥大学)
- Shanghai University of Finance and Economics(上海财经大学)
- University of Massachusetts at Lowell(马萨诸塞大学洛厄尔分校)
- The Hong Kong University of Science and Technology(香港科技大学)
- University of Illinois at Chicago(伊利诺伊大学芝加哥分校)
- Jilin University(吉林大学)
- The University of Hong Kong(香港大学)
- Michigan State University(密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。