MMAligner:通过表示校准保护多模态大语言模型
MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration
浏览论文内容
中文总结 AI 辅助
MMAligner通过校准多模态大语言模型的表示,将不安全多模态输入的拒绝率提升至99%,仅造成不足2%的效用下降,显著优化了安全与效用的权衡。
中文摘要 AI 辅助
多模态大语言模型(Multimodal Large Language Models, MLLMs)通常会拒绝不安全的文本提示,但对语义等价的多模态输入却会生成有害响应。现有防御方法要么依赖外部防护措施,这会增加推理开销且无法修复内在缺陷;要么采用安全微调,将对齐视为黑盒优化,可能牺牲效用或需要大量多模态数据集。为探究这种安全差异的原因,我们从几何角度分析MLLM的表示,发现从文本中学习到的安全机制会跨模态存在:共享的安全子空间和拒绝边界仍然有效,位于该边界内的表示会持续触发拒绝。然而,不安全的多模态输入会发生表示偏移,使其大部分位于边界之外,从而绕过模型的内在安全机制。这表明多模态安全退化源于表示不对齐,而非缺乏安全能力。基于此发现,我们提出MMAligner,一种将不安全多模态表示校准到预存拒绝区域的防御方法。MMAligner应用硬下界确保拒绝,应用软上界避免过度修改,并为良性输入设置保留目标。在多个开源MLLMs上的实验表明,MMAligner将不安全多模态输入的平均拒绝率提升至99%,同时效用下降不足2%,且仅需极少训练数据,大幅改善了安全-效用权衡关系,优于现有基线方法。
英文摘要
Multimodal large language models (MLLMs) often refuse unsafe text prompts yet generate harmful responses to semantically equivalent multimodal inputs. Existing defenses either rely on external guardrails, which add inference overhead without repairing intrinsic flaws, or safety fine-tuning, which treats alignment as black-box optimization and may sacrifice utility or require large multimodal datasets. To identify the cause of this safety disparity, we analyze MLLM representations geometrically. We find that safety mechanisms learned from text persist across modalities: a shared safety subspace and refusal boundary remain effective, and representations inside this boundary consistently trigger refusals. However, unsafe multimodal inputs undergo a representation shift that places most of them outside the boundary, allowing them to bypass the model's intrinsic safety mechanism. This indicates that multimodal safety degradation stems from representation misalignment rather than the absence of safety capability. Based on this finding, we propose MMAligner, a safeguarding method that calibrates unsafe multimodal representations into the pre-existing refusal region. MMAligner applies a hard lower bound to ensure refusal, a soft upper bound to avoid excessive modification, and a preservation objective for benign inputs. Experiments across multiple open-source MLLMs show that MMAligner raises the average refusal rate on unsafe multimodal inputs to 99% with less than 2% utility degradation and minimal training data, substantially improving the safety-utility trade-off over existing baselines. (*Due to the notification from arXiv, "The Abstract field cannot be longer than 1,920 characters", the Abstract that appeared is shortened.)