SafeStyle:用于扩散风格化中可控风格泄漏权衡的校准风格残差注入
SafeStyle: Calibrated Style Residual Injection for Controllable Style-Leakage Trade-off in Diffusion Stylization
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SafeStyle提出免训练的校准风格残差注入方法,通过子空间估计和残差范数预算,在扩散风格化中平衡风格保真度与内容泄漏,实验显示高风格相似度且低语义泄漏。
AI中文摘要:
参考引导的扩散风格化旨在从参考图像转移视觉风格,同时保留文本提示指定的语义。然而,图像条件作用常常将可转移的风格线索与参考特定的内容纠缠在一起,导致固有的权衡:更强的条件作用提高了风格保真度,但增加了内容泄漏,而激进的抑制则减少了泄漏,却以风格表达为代价。这一挑战因纹理主导和几何主导风格在空间组织上的显著差异而进一步复杂化。为解决这些问题,我们提出了SafeStyle,一种在冻结扩散模型中进行校准风格残差注入的免训练框架。SafeStyle首先从紧凑的校准集中估计风格支持和内容关联子空间,保留它们的信息重叠部分,同时抑制无用的内容变化。然后,它通过自适应空间粒度传输纯化后的风格证据,并通过显式的残差范数预算约束其有效影响。在纹理主导和几何主导风格上的实验表明,SafeStyle实现了0.432的DINO风格相似度,同时保持了有竞争力的文本对齐。在语义不相交的泄漏压力基准上,它进一步实现了0.474的DINO风格相似度,仅0.8%的语义泄漏,展示了风格保真度与参考内容抑制之间的有效平衡。
英文摘要:
Reference-guided diffusion stylization aims to transfer visual style from a reference image while preserving the semantics specified by a text prompt. However, image conditioning often entangles transferable style cues with reference-specific content, leading to an inherent trade-off: stronger conditioning improves style fidelity but increases content leakage, whereas aggressive suppression reduces leakage at the cost of style expression. This challenge is further complicated by the distinct spatial organization of texture- and geometry-dominant styles. To address these issues, we propose SafeStyle, a training-free framework for calibrated style residual injection in frozen diffusion models. SafeStyle first estimates style-supported and content-associated subspaces from compact calibration sets, preserving their informative overlap while suppressing useless content variations. It then transports the purified style evidence over adaptive spatial granularity and constrains its effective influence through an explicit residual-norm budget. Experiments across texture- and geometry-dominant styles show that SafeStyle achieves a DINO style similarity of 0.432 while maintaining competitive text alignment. On a semantically disjoint leakage-stress benchmark, it further achieves a DINO style similarity of 0.474 with only 0.8\% semantic leakage, demonstrating an effective balance between style fidelity and reference-content suppression.