arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ShieldCLIP:多模态基础模型中有害内容缓解的选择性安全对齐

ShieldCLIP: Selective Safety Alignment for Harmful Content Mitigation in Multimodal Foundation Models

Tobia Poppi, Silvia Cappelletti, Samuele Poppi, Marcella Cornia, Lorenzo Baraldi, Diego Garcia-Olano, Rita Cucchiara

arXiv 2609.39688首次发表:更新:

发表机构

University of Modena and Reggio Emilia; University of Pisa; MBZUAI; Meta Superintelligence Labs(摩德纳大学与雷焦艾米利亚大学; 比萨大学; 穆罕默德·本·扎耶德人工智能大学; Meta超级智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ShieldCLIP通过逐模态安全状态条件化对齐,引入ViSUv2数据集和四路目标,在检索与生成任务中减少有害输出并保留原始嵌入效用。

AI 中文摘要

多模态编码器(如CLIP)支撑着许多下游系统,但其网络规模训练数据中嵌入了有害关联,安全对齐必须抑制这些关联,同时不应不必要地改变良性表征。由于伦理和实际限制阻止了大规模收集真实的不安全内容,现有数据集将安全的真实样本与生成的对应样本配对,但将每个生成的样本标记为不安全,即使单个模态本身是安全的。为解决这一问题,我们引入了ShieldCLIP,这是第一个将安全对齐条件化于每个模态的观察到的安全状态而非样本来源的框架,从而在仅重定向不安全内容的同时保留安全内容。我们还引入了ViSUv2,一个包含195k四元组的数据集,具有跨578个概念和28个类别的独立逐模态安全标签。利用这些标签,ShieldCLIP定义了一个超越成对监督的四路条件目标:安全内容被锚定,不安全模态被重定向到其安全对应物,混合对仅更新不安全分支,当两者都不安全时强制执行一致性。我们在跨模态检索、使用Stable Diffusion v1.4和SDXL的文本到图像生成以及使用LLaVA的图像到文本生成上评估了ShieldCLIP。在这些设置中,与先前的安全对齐编码器和强缓解基线相比,ShieldCLIP持续减少有害输出,同时保留原始嵌入空间的效用。广泛的消融研究进一步表明,模态特定监督和选择性对齐目标都有助于这些改进。源代码、训练模型和ViSUv2(在受控访问协议下)将在此https URL公开提供。

英文摘要

Multimodal encoders such as CLIP underlie many downstream systems, but their web-scale training data embed harmful associations that safety alignment must suppress without unnecessarily changing benign representations. Because ethical and practical constraints prevent collecting real unsafe content at scale, existing datasets pair safe real samples with generated counterparts, but label every generated sample unsafe, even when one modality is individually safe. To address this, we introduce ShieldCLIP, the first framework to condition safety alignment on the observed safety state of each modality rather than the origin of a sample, preserving safe content while redirecting only what is unsafe. We also introduce ViSUv2, a 195k-quadruplet dataset with independent per-modality safety labels across 578 concepts and 28 categories. Using these labels, ShieldCLIP defines a four-way conditional objective beyond pair-level supervision: safe content is anchored, unsafe modalities are redirected to their safe counterparts, mixed pairs update only the unsafe branch, and coherence is enforced when both are unsafe. We evaluate ShieldCLIP on cross-modal retrieval, text-to-image generation with Stable Diffusion v1.4 and SDXL, and image-to-text generation with LLaVA. Across these settings, ShieldCLIP consistently reduces harmful outputs over prior safety-aligned encoders and strong mitigation baselines, while preserving the utility of the original embedding space. Extensive ablation studies further show that both modality-specific supervision and the selective alignment objective contribute to these gains. Source code, trained models, and ViSUv2 (under a controlled-access protocol) will be made publicly available at https://aimagelab.github.io/ShieldCLIP/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑