SafeNexus:在多模态大语言模型(MLLMs)中发现与调控模态通用安全神经元
SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs
浏览论文内容
中文总结 AI 辅助
SafeNexus是一种跨模态安全对齐框架,通过定位并调控模态通用安全神经元,提升多模态大语言模型在跨模态威胁下的安全性,且能保留模型效用。
中文摘要 AI 辅助
尽管大语言模型(LLMs)已展现出良好的安全性能,但将其扩展至多模态大语言模型(MLLMs)时,会暴露出多模态能力扩展与现有安全机制之间的显著差距。当前的防御措施大多局限于特定模态场景,从而限制了其应对更广泛跨模态威胁的鲁棒性。为弥合这一差距,我们提出SafeNexus,这是一种采用专用神经元级干预策略的跨模态安全对齐框架。首先,我们构建了一种神经元定位范式,通过刻画中间层激活模式识别功能专用神经元,并通过重要性评分量化其功能显著性。基于该范式,我们利用对比数据识别模态绑定安全神经元(BS-Neurons),并通过针对性抑制验证其在各模态内调控安全行为的作用。进一步的跨模态分析将模态通用安全神经元(US-Neurons)定义为在各模态中识别出的BS-Neurons的共享子集,作为防御有害跨模态攻击的核心。我们观察到,抑制这些神经元会显著降低各模态的安全性能,同时几乎不影响整体效用。基于这些发现,我们提出两种安全对齐策略:激活级安全放大器和安全神经元校准器。所提策略通过两条不同路径提升模型安全性:前者放大US-Neurons的激活幅度,后者通过针对性微调选择性校准US-Neurons。大量实验表明,我们的方法在涵盖不同模态组合的安全基准上优于现有最先进方法,同时有效保留了模型效用。
英文摘要
Although Large Language Models (LLMs) have demonstrated promising safety performance, extending them to Multimodal Large Language Models (MLLMs) exposes a significant gap between expanded multimodal capabilities and existing safety mechanisms. Current defenses remain predominantly confined to specific modal settings, thereby limiting their robustness against broader cross-modal threats. To bridge this gap, we introduce SafeNexus, a cross-modal safety alignment framework that adopts a dedicated neuron-level intervention strategy. First, we formulate a neuron localization paradigm that identifies functionally specialized neurons by characterizing intermediate-layer activation patterns and quantifying their functional salience through importance scoring. Building upon this paradigm, we exploit contrastive data to identify modality-bound safety neurons (BS-Neurons), and validate their role in regulating safety behavior within each modality via targeted suppression. Further cross-modal analysis defines modality-universal safety neurons (US-Neurons) as the shared subset of BS-Neurons identified across individual modalities, serving as the core for defending against harmful cross-modal attacks. We observe that suppressing these neurons substantially degrades safety performance across modalities, while leaving overall utility largely unaffected. Building on these insights, we propose two safety alignment strategies: activation-level safety amplifier and safety neuron calibrator. The proposed strategies enhance model safety through two distinct routes: the former amplifies the activation magnitudes of US-Neurons, while the latter selectively calibrates them via targeted fine-tuning. Extensive experiments demonstrate that our method outperforms prevailing state-of-the-art approaches on safety benchmarks spanning diverse modality combinations, while effectively preserving utility.