发表机构
University of Science, Ho Chi Minh City; Vietnam National University, Ho Chi Minh City(胡志明市理科大学; 越南国立大学胡志明市分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出免训练扩散风格迁移框架TRACE,通过内容-风格分解和时间自适应残差注意力控制,在提升风格保真度的同时减少内容泄漏,实验显示显著优于基线。
AI 中文摘要
参考引导的风格迁移旨在保留内容图像的语义结构,同时迁移风格参考的视觉外观。近期基于扩散的方法通过利用强大的预训练生成先验,实现了令人印象深刻的风格化质量。然而,免训练方法仍面临风格保真度、内容保留和内容泄漏之间的艰难权衡。直接注入风格可能无意中迁移风格图像的语义内容,而固定的引导调度往往忽略扩散采样的时间和状态依赖性。为解决这些局限,我们提出TRACE,一种具有时间自适应残差注意力控制和内容-风格分解的免训练扩散风格迁移框架。TRACE首先进行离线CLIP子空间分析,从配对数据中分离内容和风格方向。在推理期间,它从风格参考中移除内容相关组件,并从内容参考中移除风格相关组件以减少泄漏。然后,它通过残差交叉注意力注入风格信息,并应用不确定性感知引导,在每个去噪步骤自适应引导信号。实验表明,TRACE在风格化和保留之间实现了有利的权衡。与基于最优控制的基线相比,TRACE显著提高了风格保真度(+17.28 CSD和+34.10 SRA)。同时,与风格化方法相比,它更好地保留了内容结构(+12.80 DINO,+5.52 CLIP-I,和-8.19 LPIPS),并将DCL中的方向性语义泄漏减少了29.5%。我们的代码在此https URL公开可用。
英文摘要
Reference-guided style transfer aims to preserve the semantic structure of a content image while transferring the visual appearance of a style reference. Recent diffusion-based methods achieve impressive stylization quality by exploiting strong pretrained generative priors. However, training-free approaches still face a difficult trade-off among style fidelity, content preservation, and content leakage. Direct style injection may unintentionally transfer semantic content from the style image, while fixed guidance schedules often ignore the time- and state-dependent nature of diffusion sampling. To address these limitations, we propose TRACE, a training-free diffusion style transfer framework with Time-adaptive Residual Attention Control and Content-Style Decomposition. TRACE first performs offline CLIP-based subspace analysis to separate content and style directions from paired data. During inference, it removes content-related components from the style reference and style-related components from the content reference to reduce leakage. It then injects style information through residual cross-attention and applies uncertainty-aware guidance to adapt the guidance signal at each denoising step. Experiments show that TRACE achieves a favorable trade-off between stylization and preservation. Compared with optimal-control-based baselines, TRACE substantially improves style fidelity (+17.28 CSD and +34.10 SRA). While, compared with stylization methods, it better preserves content structure (+12.80 DINO, +5.52 CLIP-I, and -8.19 LPIPS) and reduces directional semantic leakage by 29.5% in DCL. Our code is publicly available at https://github.com/pixelchemy-research/TRACE.
CommentsAccepted to ACCV 2026