发表机构
Beihang University; China CITIC Bank; Stable AI; Tsinghua University(北京航空航天大学; 中国中信银行; Stable AI; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SafeMol提出双模态安全对齐框架,构建SafeMolBench基准,联合优化轻量模块并显式建模分子危险性与有害意图,显著降低越狱攻击成功率并保持任务实用性。
AI 中文摘要
分子多模态模型支持多样化的理解和生成任务,但在处理危险分子时可能引入安全漏洞。在本工作中,我们揭示了在纯文本和图条件设置下的实质性越狱漏洞。我们的分析进一步表明,安全鲁棒性必须在所有输入模态中保持一致,同时平衡安全性、过度拒答和实用性。为应对这些挑战,我们构建了SafeMolBench,一个包含3702个样本、覆盖618种独特危险分子和安全的分子任务的分子多模态安全对齐基准,将其组织为危险-有害、危险-允许和实用性回放三个子集,以支持安全、过度拒答和实用性的统一训练与评估。基于SafeMolBench,我们提出了SafeMol,一个参数高效的安全对齐框架,该框架在纯文本和图条件输入上联合优化轻量级模块,使用MMD进行分布级表示对齐以减少模态引起的差异,并显式建模分子的危险性和有害操作意图。在SafeMolBench上的实验表明,SafeMol将攻击成功率降低了数十个百分点,同时基本保持了较低的过度拒答率并保留了分子任务的实用性。
英文摘要
Molecular multimodal models support diverse understanding and generation tasks but may introduce safety vulnerabilities when handling hazardous molecules. In this work, We reveal substantial jailbreak vulnerabilities under both text-only and graph-conditioned settings. Our analysis further shows that safety robustness must hold across input modalities while balancing safety, over-refusal, and utility. To address these challenges, we construct SafeMolBench, a molecular multimodal safety-alignment benchmark with 3702 samples covering 618 unique hazardous molecules and safe molecular tasks, organized into hazardous-harmful, hazardous-allowed, and utility-replay subsets to support unified training and evaluation of safety, over-refusal, and utility. Based on SafeMolBench, we propose SafeMol, a parameter-efficient safety alignment framework that jointly optimizes lightweight modules across text-only and graph-conditioned inputs, uses MMD for distribution-level representation alignment to reduce modality-induced discrepancies, and explicitly models molecular hazardousness and harmful operational intent. Experiments on SafeMolBench show that SafeMol reduces attack success by several tens of percentage points while largely maintaining low over-refusal and preserving molecular-task utility.