arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19719cs.CVcs.AI

用于无风格编码器的扩散风格化的尺度分离条件设置

Scale-Separated Conditioning for Style-Encoder-Free Diffusion Stylization

  • College of Computing, Georgia Institute of Technology(佐治亚理工学院计算学院)
  • Courant Institute of Mathematical Sciences, New York University(纽约大学柯朗数学科学研究所)
  • Department of Computer Science, Brown University(布朗大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

Jingtao Zhang, Haorui Gao, Youqing Liang, Zeming Liu

AI总结:

本研究提出无风格编码器的扩散风格化框架SEFS,通过单张图像低分辨率裁剪形成风格令牌,在未配对单张图像上训练,提升艺术风格化的内容一致性等性能,代码将公开。

AI中文摘要:

基于参考的扩散风格化需要将目标几何与可迁移外观分离。现有基于调优的方法通常依赖对齐的内容-风格-目标三元组或辅助视觉编码器,这增加了数据成本,且可能从风格参考中迁移意外的场景结构。我们提出SEFS(Style-Encoder-Free Stylization,无风格编码器的风格化),一种用于扩散Transformer的无风格编码器的条件设置框架。SEFS从单张训练图像的随机低分辨率裁剪中形成风格令牌。这种裁剪瓶颈保留了调色板、笔触、纹理和材质等局部外观统计,同时减少了对全局布局线索的访问。目标内容通过边缘和分割线索编码,并通过参数高效的可训练投影与噪声潜变量融合。我们添加风格到去噪的重归一化以实现令牌统计对齐,以及跨块跳跃融合以实现空间细节。SEFS在未配对的单张图像上训练;冻结的扩散VAE仅用于将图像条件放置在潜空间中。在艺术风格化基准上,SEFS在保留参考风格亲和力的同时,提高了内容一致性和泄漏诊断能力,且 ablation 实验支持裁剪分辨率、重归一化和跳跃融合的选择。SEFS的代码将公开提供。

英文摘要:

Reference-based diffusion stylization requires separating target geometry from transferable appearance. Existing tuning-based methods often rely on aligned content-style-target triplets or auxiliary visual encoders, which increases data cost and can transfer unintended scene structure from the style reference. We propose SEFS (Style-Encoder-Free Stylization), a style-encoder-free conditioning framework for diffusion transformers. SEFS forms style tokens from stochastic low-resolution crops of single training images. This crop bottleneck preserves local appearance statistics such as palette, stroke, texture, and material, while reducing access to global layout cues. Target content is encoded by edge and segmentation cues and fused with the noisy latent through parameter-efficient trainable projections. We add style-to-denoising re-normalization for token-statistic alignment and cross-block skip fusion for spatial detail. SEFS trains on unpaired single images; the frozen diffusion VAE is used only to place image conditions in the latent space. On artistic stylization benchmarks, SEFS improves content consistency and leakage diagnostics while retaining reference-style affinity, and ablations support the crop-resolution, re-normalization, and skip-fusion choices. The code of SEFS will be made publicly available.

↑