arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08517cs.MM

扩散基础模型治理中的概念级风险与校准

Concept-Level Risk and Calibration for Governance in Diffusion Foundation Models

  • Nanjing University of Aeronautics and Astronautics(南京航空航天大学)
  • Jiangxi University of Finance and Economics(江西财经大学)
  • City University of Hong Kong(香港城市大学)
  • University of Insubria(因苏布里亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Kun Xu, Yushu Zhang, Tao Wang, Shuren Qi, Barbara Carminati, Elena Ferrari, Yuming Fang

AI总结:

本文提出概念级概率审计框架CLRC,通过风险算子和校准方法统一评估扩散模型在多种条件下的治理风险,实验表明嵌入访问和混淆提示风险常被低估。

AI中文摘要:

扩散模型已成为多媒体生成的核心范式,为个性化、语义编辑和选择性遗忘提供了强大的概念驱动可控性。然而,随着语义控制从自然语言提示扩展到学习嵌入和干预管道,这些系统的安全性和治理变得日益难以统一评估,尤其是对于安全敏感、身份关联及其他隐私相关概念。现有研究主要依赖启发式审计、对抗性探测或特定任务的擦除基准,因此对跨模型、条件通道和部署条件的系统比较支持有限。我们提出了一种用于扩散模型的概念级概率审计与报告框架。我们将治理相关的概念行为形式化为由随机生成引发的伯努利语义事件,并定义了一个概念风险算子,将模型-通道配置映射到结构化风险概况,从而支持跨提示接口、学习嵌入通道、模型和记录条件的比较。我们应用样本级事后校准和配置级风险聚合,并表明概率误差可以改变政策边界附近的阈值化操作。在SD1.5、SD2.1和SDXL上的实验揭示了跨概念族、通道、记录条件和偏移协议的一致但非均匀的操作风险模式。特别是,基于嵌入的访问和混淆提示暴露了标准提示评估通常低估的风险。一个池化的多协议校准器提高了保留概率的可靠性,但我们不声称从仅标准校准器进行迁移。CLRC为多媒体生成系统的概率性和决策感知治理提供了通用审计模式。

英文摘要:

Diffusion models have become a core paradigm for multimedia generation, offering powerful concept-driven controllability for personalization, semantic editing, and selective unlearning. However, as semantic control extends beyond natural-language prompts to learned embeddings and intervention pipelines, the safety and governance of these systems become increasingly difficult to evaluate in a unified manner, especially for safety-sensitive, identity-linked, and other privacy-relevant concepts. Existing studies mainly rely on heuristic audits, adversarial probing, or task-specific erasure benchmarks, and therefore provide limited support for systematic comparison across models, conditioning channels, and deployment conditions. We present a concept-level probabilistic audit and reporting framework for diffusion models. We formalize governance-relevant concept behaviors as Bernoulli semantic events induced by stochastic generation, and define a Concept Risk Operator that maps model-channel configurations to structured risk profiles, enabling comparison across prompting interfaces, learned embedding channels, models, and recorded conditions. We apply sample-level post-hoc calibration and configuration-level risk aggregation, and show that probability error can change thresholded actions near policy boundaries. Experiments on SD1.5, SD2.1, and SDXL reveal consistent yet non-uniform operational risk patterns across concept families, channels, recorded conditions, and shifted protocols. In particular, embedding-based access and obfuscated prompts expose risks often understated by standard-prompt evaluation. A pooled multi-protocol calibrator improves held-out probability reliability, but we do not claim transfer from a standard-only calibrator. CLRC provides a common audit schema for probabilistic and decision-aware governance of multimedia generation systems.

补充信息

↑