发表机构
The Pennsylvania State University; University of Sheffield(宾夕法尼亚州立大学; 谢菲尔德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现有MLLM遗忘基准测试的局限性,提出PPE-Bench基准测试,包含公私纠缠图像,引入两种方法在遗忘中保护公共信息,实验发现现有方法能减少隐私泄露,但常损害相邻公共信息。
AI 中文摘要
多模态大语言模型(MLLMs)虽能力强,但可能记忆网络数据中的私人信息,引发隐私担忧。机器遗忘提供了一种无需从头重新训练即可去除此类私人知识的方法。然而,现有的MLLM遗忘基准测试有两个主要局限性。首先,它们依赖于仅包含单个目标个体的简化图像,无法反映现实世界照片的视觉复杂性。其次,它们通常假设遗忘集和保留集是完全分开的,而忽略了私人信息往往在视觉上与良性公共信息纠缠在一起的事实。例如,一个私人个体可能与公众人物一起出现或出现在著名地标前,在这种情况下,遗忘私人目标不应损害公共背景。为了解决这些局限性,我们提出了PPE-Bench,这是一个用于评估公私纠缠下MLLM遗忘能力的新基准测试。每个图像都包含一个要遗忘的目标个体和要保留的公共信息,包括公众人物和地标。我们进一步引入了两种简单但有效的方法,以便在遗忘过程中更好地保留公共信息。通过实验,我们发现现有的遗忘方法可以减少私人信息泄露,但往往会对相邻的公共信息造成实质性损害。
英文摘要
Multimodal Large Language Models (MLLMs) have shown strong capabilities, but they may memorize private information from web data, raising privacy concerns. Machine unlearning offers a way to remove such private knowledge without retraining from scratch. However, existing MLLM unlearning benchmarks have two major limitations. First, they rely on simplified images that contain only the single target individual, failing to reflect the visual complexity of real-world photos. Second, they typically assume that the forget set and retain set are fully separated, ignoring the fact that private information is often visually entangled with benign public information. For example, a private individual may appear with a public figure or in front of a well-known landmark, where unlearning the private target should not damage the public context. To address these limitations, we propose PPE-Bench, a new benchmark for evaluating MLLM unlearning under private-public entanglement. Each image contains a target individual to be forgotten and public information to be preserved, including public figure and landmark. We further introduce two simple but effective methods to better preserve public information during unlearning. Through experiments, we find that existing unlearning methods can reduce private information leakage, but often substantially harm adjacent public information.
Commentsto appear in EMNLP 2026