发表机构
University of Trento; Fondazione Bruno Kessler(特伦托大学; 布鲁诺·凯斯勒基金会)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多模态大语言模型遗忘中的公平性问题,提出FAIRGET基准和FAUN算法,通过模拟现实场景下不平衡的遗忘请求,在考虑数据不平衡性质时遗忘身份,实验证明该方法在遗忘质量和公平性上具有优越性。
AI 中文摘要
机器遗忘已成为从训练模型中删除个人数据以符合近期人工智能法规的一种工具。为评估多模态大语言模型(MLLM)中的遗忘有效性,先前工作在虚拟身份上微调模型,模拟对这些身份子集的遗忘请求,这些请求通常均匀分布。但在现实场景中,不同人口群体的遗忘请求频率不同,可能改变模型对这些群体的内部信念并导致偏差行为。为填补这一空白,我们提出FAIRGET,首个在不平衡、现实的遗忘请求下评估遗忘的视觉问答基准。这些请求模拟多种现实场景,从简单到具有挑战性,若不考虑公平性会导致有偏差的遗忘模型。此外,我们提出FAUN,首个用于MLLM的遗忘算法,在保留模型公平性的同时遗忘数据。FAUN利用偏差感知激活引导机制在考虑遗忘数据不平衡性质的情况下遗忘身份。在FAIRGET和已有的FIUBench上的实验证明了我们方法在遗忘质量和公平性上的优越性。
英文摘要
Machine unlearning has emerged as a tool for removing personal data from trained models to comply with recent AI regulations. To evaluate unlearning effectiveness in multimodal large language models (MLLMs), prior works fine-tune models on fictitious identities, simulating unlearning requests on subsets of these IDs, which are typically uniformly distributed. However, in realistic scenarios, people from different demographic groups may request to be unlearned at different frequencies, potentially altering the model's internal beliefs for these groups and leading to biased behaviors. To fill this gap, we propose FAIRGET, the first Visual Question Answering benchmark that evaluates unlearning under unbalanced, realistic, forget requests. These requests are designed to simulate multiple realistic scenarios, ranging from simple to challenging settings, that lead to biased unlearned models if fairness is not accounted for. Additionally, we propose FAUN, the first unlearning algorithm for MLLMs that forgets unlearning data while preserving model fairness. FAUN exploits a bias-aware activation steering mechanism to unlearn identities while accounting for the unbalanced nature of the forget data. Experiments on FAIRGET and the established FIUBench demonstrate our method's superiority both in unlearning quality and fairness.
Comments33 pages