arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2511.19558cs.CRcs.AIcs.CVcs.LG

SPQR:良性模型适应下安全对齐的多维基准测试

SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation

  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • University of Waterloo(滑铁卢大学)
  • Michigan State University(密歇根州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Mohammed Talha Alam, Nada Saadi, Fahad Shamshad, Nils Lukas, Karthik Nandakumar, Fahkri Karray, Samuele Poppi

AI总结:

研究文本到图像扩散模型安全对齐在良性微调下的稳定性,引入SPQR基准测试,通过单评分指标提供统一框架,经多方面分析确定安全对齐失败情况,为T2I安全对齐技术提供简洁全面的基准。

AI中文摘要:

文本到图像的扩散模型可能会生成受版权保护、不安全或私密的内容。安全对齐旨在抑制特定概念,但评估很少测试在部署后常规应用的良性下游微调(例如LoRA个性化、风格/域适配器)下安全性是否持续。我们研究了当前安全方法在良性微调下的稳定性,发现频繁出现故障。由于真正的安全对齐必须经受住即使是良性的部署后适应,我们引入了SPQR基准测试(安全、提示遵循、质量和鲁棒性)。SPQR是一个单评分指标,通过报告单个排行榜分数来促进比较,提供一个统一、可重复的框架,以评估安全对齐的扩散模型在良性微调下如何很好地保持安全性、实用性和鲁棒性。我们进行了多语言、特定领域和分布外分析,以及按类别细分,以确定良性微调后安全对齐何时失败,最终展示了SPQR作为T2I模型的T2I安全对齐技术的简洁而全面的基准测试。

英文摘要:

Text-to-image diffusion models can emit copyrighted, unsafe, or private content. Safety alignment aims to suppress specific concepts, yet evaluations seldom test whether safety persists under benign downstream fine-tuning routinely applied after deployment (e.g., LoRA personalization, style/domain adapters). We study the stability of current safety methods under benign fine-tuning and observe frequent breakdowns. As true safety alignment must withstand even benign post-deployment adaptations, we introduce the SPQR benchmark (Safety, Prompt adherence, Quality, and Robustness). SPQR is a single-scored metric that provides a unified, reproducible framework to evaluate how well safety-aligned diffusion models preserve safety, utility, and robustness under benign fine-tuning, by reporting a single leaderboard score to facilitate comparisons. We conduct multilingual, domain-specific, and out-of-distribution analyses, along with category-wise breakdowns, to identify when safety alignment fails after benign fine-tuning, ultimately showcasing SPQR as a concise yet comprehensive benchmark for T2I safety alignment techniques for T2I models.

补充信息

↑