arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CLeaR:解决风格迁移中泄漏-退化困境的统一框架

CLeaR: A Unified Framework for Resolving the Leakage-Degradation Dilemma in Style Transfer

Teng Zhou, Yunhao Chen

arXiv 2609.38136首次发表:更新:

发表机构

Zhejiang University; Fudan University(浙江大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CLeaR提出无训练框架,通过正交子空间投影、集成反演和能量引导校准,解决风格迁移中内容泄漏与风格保真度的困境,在StyleBench上改善风格对齐并减少泄漏。

AI 中文摘要

风格迁移旨在以参考图像的风格渲染目标内容,但现有方法常常遭受内容泄漏问题,即风格参考中的物体、布局或语义出现在生成输出中。尽管先前的数据驱动和无训练方法可以减少泄漏,它们往往面临泄漏-退化困境:更强的内容抑制可能削弱风格保真度,而更丰富的风格保留可能重新引入不需要的参考内容。我们在完整的风格迁移流程中识别了这一困境,包括特征分离、特征空间接地和扩散生成。为解决这些问题,我们提出了CLeaR,一个用于抗内容泄漏风格迁移的无训练框架。CLeaR首先使用正交子空间投影在每个视觉基础模型(VFM)特征空间中定义内容减少的风格目标。然后执行集成反演,优化一个共享的像素空间风格锚点,以满足多个VFM的风格约束。最后,能量引导校准通过将去噪轨迹导向集成定义的风格流形,在扩散采样期间保持风格对齐。我们进一步提供了理论分析,表明风格锚点估计误差随着VFM数量的增加而减小。在StyleBench上的实验表明,与现有方法相比,CLeaR改善了风格对齐,减少了内容泄漏,并取得了更好的LLM-as-Judge评估结果。代码可在\ref{this https URL}{this https URL}获取。

英文摘要

Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and training-free methods can reduce leakage, they often face a leakage-degradation dilemma: stronger content suppression may weaken style fidelity, while richer style preservation may reintroduce unwanted reference content. We identify this dilemma across the full style-transfer pipeline, including feature separation, feature-space grounding, and diffusion generation. To address these issues, we propose CLeaR, a training-free framework for content-leakage-resistant style transfer. CLeaR first uses Orthogonal Subspace Projection to define content-reduced style targets in each vision foundation model (VFM) feature space. It then performs Ensemble Inversion, which optimizes a shared pixel-space style anchor satisfying style constraints across multiple VFMs. Finally, Energy-Guided Calibration maintains style alignment during diffusion sampling by steering the denoising trajectory toward the ensemble-defined style manifold. We further provide a theoretical analysis showing that the style-anchor estimation error decreases with the number of VFMs. Experiments on StyleBench demonstrate that CLeaR improves style alignment, reduces content leakage, and achieves better LLM-as-Judge evaluation compared with existing methods. The code is available at \href{https://github.com/0606zt/CLeaR}{https://github.com/0606zt/CLeaR}.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑