发表机构
Emory University(埃默里大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对持续学习下的多模态大语言模型,研究人员提出持久公平性后门攻击,通过两种机制注入持久群体歧视,该攻击能规避标准防御且在多轮持续学习中留存。
AI 中文摘要
多模态大语言模型(MLLMs)正越来越多地被部署到公平性作为关键安全要求的高风险领域。在实际应用中,这些模型会通过持续学习(CL)不断更新,以适应不断变化的任务和数据分布。已有研究表明,后门攻击可通过隐藏触发器操纵MLLM的响应,但在模型经历后续的CL更新时,朴素植入的后门会失效。尽管公平性已成为MLLM部署的核心关注点,但后门诱导的公平性违反是否能在CL中留存尚未得到探索,这留下两个关键问题未解答:(1)后门能否可靠地在MLLMs中诱导公平性违反;(2)此类针对公平性的后门能否在持续学习中持久存在。我们通过提出持久公平性后门攻击(PFBA)来弥合这一差距,以向MLLMs中注入持久且针对特定群体的歧视。具体而言,PFBA通过两种新机制实现这一目标:潜在空间公平性强化通过将特权群体的表示锚定以保留效用,同时排斥并聚类目标群体的表示以维持歧视,从而重塑模型的深层特征几何;持续学习模拟通过针对模拟的参数漂移迭代优化触发器,以确保后门在未来的更新中持久存在。大量实验表明,PFBA会诱导严重的公平性差异,且这些差异会在多轮持续学习中持续存在,同时能规避标准后门防御。相关数据和代码可在此https URL获取。
英文摘要
Multimodal Large Language Models (MLLMs) are increasingly deployed in high-stakes domains where fairness is a critical safety requirement. In practice, these models are continually updated through continual learning (CL) to adapt to evolving tasks and data distributions. Prior work has shown that backdoor attacks can manipulate MLLM responses through hidden triggers, but naively implanted backdoors degrade as models undergo subsequent updates of CL. Although fairness has emerged as a central concern for MLLM deployment, whether backdoor-induced fairness violations can survive CL remains unexplored, leaving two critical questions unanswered: (1) whether a backdoor can reliably induce fairness violations in MLLMs, and (2) whether such fairness-targeted backdoors can persist through continual learning. We bridge this gap by proposing Persistent Fairness Backdoor Attack (PFBA) to inject persistent and group-specific discrimination into MLLMs. Specifically, PFBA achieves this through two novel mechanisms. The Latent Space Fairness Reinforcement reshapes the model's deep feature geometry by anchoring privileged-group representations to preserve utility while repelling and clustering targeted-group representations to sustain discrimination, and the Continual Learning Simulation iteratively optimizes the trigger against simulated parameter drift to ensure backdoor persistence across future updates. Extensive experiments demonstrate that PFBA induces severe fairness disparities that persist across continual learning rounds, evading standard backdoor defenses. The data and code are publicly available at https://github.com/lyygua/PFBA.
CommentsCIKM 2026