arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

统一多语言多模态的LVLM安全对齐:一个通用锚点

One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs

Enyi Shi, Fei Shen, Chuancheng Shi, Linxia Zhu, Shuyi Miao, Jinhui Tang, Tat-Seng Chua

arXiv 2607.27917首次发表:更新:

AI 中文总结

针对LVLM面临的多语言多模态复合攻击防御难题,提出MLS-Neurons驱动的跨维度安全对齐框架,仅更新约0.03%参数实现安全监督迁移,在多语言多模态安全基准上性能优于SOTA且保留通用实用性。

AI 中文摘要

随着大型视觉语言模型(LVLM)在全球部署,多语言指令与视觉信息的结合使得恶意攻击比以往更隐蔽、更复杂。然而,现有方法将语言与模态防御隔离,加之安全数据稀缺、微调成本高,导致模型难以抵御复合攻击。为应对这一严峻挑战,我们提出一种由模态与语言共享安全神经元(MLS-Neurons)驱动的神经元级跨维度安全对齐框架。首先,通过对比有害样本与良性样本的响应,识别单语言单模态安全神经元,利用激活强度与下游影响量化功能显著性。接着,在每种语言内对这些单模态神经元取交集,提取对视觉和文本风险均有响应的模态共享安全神经元(MS-Neurons),弥合模态间的安全表征差距。此外,以英语为语义锚点,对跨语言的MS-Neurons取交集,识别出模态与语言共享安全神经元(MLS-Neurons),作为抵御复合攻击的关键防御手段。最后,仅更新这一极小的共享神经元子集(约占参数的0.03%),将仅针对英语的安全监督迁移至多语言多模态场景。大量实验表明,我们的方法在各类多语言多模态安全基准上显著优于现有最优方法,同时保留了通用实用性。

英文摘要

As large vision-language models (LVLMs) are deployed globally, the combination of multilingual instructions and visual information makes malicious attacks more covert and sophisticated than ever before. However, existing methods isolate language and modality defenses, which, coupled with the scarcity of safety data and high fine-tuning costs, makes it difficult for models to defend against compound attacks. To address this severe challenge, we propose a neuron-level cross-dimensional safety alignment framework driven by modality- and language-shared safety neurons (MLS-Neurons). First, we identify monolingual and unimodal safety neurons by comparing responses to harmful and benign samples, quantifying functional saliency through activation strength and downstream impact. Then, by intersecting these unimodal neurons within each language, we extract modality-shared safety neurons (MS-Neurons) responsive to both visual and textual risks, bridging the safety representation gap between modalities. Furthermore, using English as a semantic anchor, we intersect MS-Neurons across languages to identify modality- and language-shared safety neurons (MLS-Neurons), serving as key defenses against compound attacks. Finally, we update only this minimal subset of shared neurons (~0.03% of parameters), transferring English-only safety supervision to multilingual and multimodal scenarios. Extensive experiments show that our method significantly outperforms state-of-the-art approaches across diverse multilingual and multimodal safety benchmarks while preserving general utility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑