arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

随机元遗忘:连接语言主干与多模态遗忘

Stochastic Meta-Unlearning: Bridging Language Backbone and Multimodal Unlearning

Zijie Liu, Jinhao Duan, Gaowen Liu, Sijia Liu, Tianlong Chen

arXiv 2607.18615首次发表:更新:

发表机构

UNC at Chapel Hill; Cisco Research; Michigan State University(北卡罗来纳大学教堂山分校; 思科研究院; 密歇根州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言模型的机器遗忘问题,提出随机元遗忘框架,利用VLM级反馈学习初始化,经内、外循环更新语言主干,实验证明该方法在遗忘-保留权衡上表现最佳,能提升准确率且可迁移。

AI 中文摘要

视觉语言模型(VLM)的机器遗忘研究仍未充分探索。与语言模型不同,VLM将语言主干与视觉组件结合,使遗忘更复杂。从单模态遗忘到VLM遗忘存在惊人现象:独立语言主干遗忘的目标在完整VLM给出图像信息时仍可恢复,这表明仅文本反馈不足以实现可靠的VLM遗忘。基于此,提出随机元遗忘(SMU),这是一个双层框架,利用VLM级反馈学习可用于遗忘的初始化。内循环中,SMU使用文本数据对语言主干进行一些遗忘步骤。外循环中,SMU将更新后的主干与冻结的VLM重新组合,并在VLM级别评估遗忘和效用。实验表明SMU实现了最佳的整体遗忘-保留权衡,相比最强基线,平均遗忘准确率降低1…

英文摘要

Machine unlearning for vision-language models (VLMs) remains underexplored. Unlike language models, VLMs combine a language backbone with visual components, which makes unlearning more complex. There is a surprising phenomenon when moving from single-modality unlearning to VLM unlearning: a target forgotten by the standalone language backbone can still be recovered when image information is given to the full VLM. This shows that text-only feedback is not enough for reliable VLM unlearning. Motivated by this observation, we propose Stochastic Meta-Unlearning (SMU), a bilevel framework that uses VLM-level feedback to learn an unlearning-ready initialization. In the inner loop, SMU applies a few unlearning steps to the language backbone using text data. In the outer loop, SMU recomposes the updated backbone with the frozen VLM and evaluates forgetting and utility at the VLM level. This design makes the unlearning update aware of the final multimodal behavior, while still keeping the update local to the language backbone. Experiments on two VLMs, two multimodal meme datasets, and three baselines show that SMU achieves the best overall forget-retain trade-off. Compared with the strongest baseline for each metric, SMU reduces average Forget accuracy by 10.52 points and improves average Retain and Test accuracy by 20.10 and 17.01 points, respectively. More importantly, SMU also transfers to new forgetting targets and to different meta-test unlearning methods. These results suggest that VLM-level feedback can make language-backbone unlearning more reliable and more transferable for VLMs.

Comments15 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑