arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非所有 token 都平等:多模态大语言模型(MLLMs)中后门的区域感知一致性修复

Not All Tokens Are Equal: Region-Aware Consistency Repair of Backdoors in MLLMs

Jiali Wei, Ming Fan, Mingkun Zhang, Haoyu Wang, Jun Sun, Guoheng Sun, Xiaoning Ren, Haijun Wang, Ting Liu

arXiv 2608.24354首次发表:更新:

发表机构

Singapore Management University(新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对 MLLMs 的后门风险,提出区域感知一致性修复框架 RACER,仅需 100 个干净样本即可有效抑制后门,平均 ASR 降至 1.1%且多数场景达 0%,同时保留模型正常效用。

AI 中文摘要

多模态大语言模型(MLLMs)正越来越多地部署在面向用户的应用中,但它们从构建所用的流水线中继承了后门风险:触发器可能存在于图像、文本或两者中。现有的模型级后门移除方法大多是为传统分类器设计的,在 MLLMs 上的有效性有限,而 MLLM 专用防御方法主要在推理时运行,过滤可疑输入却无法移除嵌入模型中的后门。为解决这一差距并从根源上消除 MLLMs 中的潜在后门,我们提出了 RACER,这是一个模型级修复框架,其核心观察是:后门会在内部表征中诱发异常的逐层演化,我们将其称为逐层不一致异常。重要的是,这种异常具有模态依赖性,主要集中在后门模型实际依赖的、编码触发器特征的 token 区域。因此,RACER 将融合表征分解为视觉和文本 token 区域,分别归一化它们的逐层不一致性,并在深层窗口上使用模态感知权重重新组合它们,从而得到一个区域感知的不一致目标,该目标能更好地捕捉局部后门诱发的异常。通过最小-最大优化,该目标驱动最坏情况扰动合成和针对所得扰动的对抗微调以修复模型,抑制后门行为所依赖的深层表征方向偏移。RACER 仅需 100 个干净样本,且无需了解触发器、攻击目标,甚至无需知道输入模型是否包含后门。在三个开源 MLLMs 上针对涵盖图像、文本和多模态触发器的 36 种后门设置进行的评估显示,RACER 将平均攻击成功率(ASR)降至 1.1%,在 32 种设置中达到 0%,同时保留了后门模型和干净模型的干净任务效用。

英文摘要

MLLMs are increasingly deployed in user-facing applications, yet they inherit backdoor risks from the pipelines used to construct them: triggers may reside in images, texts, or both. Existing model-level backdoor removal methods, largely designed for conventional classifiers, show limited effectiveness on MLLMs, while MLLM-specific defenses mainly operate at inference time, filtering suspicious inputs without removing the backdoor embedded in the model. To address this gap and eliminate latent backdoors from MLLMs at their source, we present RACER, a model-level repair framework motivated by a key observation: backdoors induce abnormal layer-to-layer evolution in internal representations, which we term the layer-wise inconsistency anomaly. Importantly, this anomaly is modality-dependent, concentrating primarily in the token region encoding the trigger features that the backdoor model actually relies on. RACER therefore decomposes the fused representation into visual and textual token regions, normalizes their layer-wise inconsistency separately, and recomposes them using modality-aware weights over a deep-layer window, yielding a region-aware inconsistency objective that better captures localized backdoor-induced anomalies. Through a min-max optimization, this objective drives worst-case perturbation synthesis and adversarial fine-tuning against the resulting perturbation to repair the model, suppressing the deep representational directional shifts on which backdoor behaviors rely. RACER requires only 100 clean samples and no knowledge of the trigger, attack objective, or even whether the input model contains a backdoor. Evaluations on three open-source MLLMs across 36 backdoor settings spanning image, text, and multimodal triggers show that RACER reduces the average ASR to 1.1%, reaching 0% in 32 settings, while preserving clean-task utility on both backdoor and clean models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑