arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

透过梵高的眼睛:基于扩散模型的全局风格迁移

Through Van Gogh's Eyes: Global Style Transfer with Diffusion Model

Jeongha Lee, Yujin Kim, Ghazanfar Ali, Suhyun Kim, Jae-In Hwang

arXiv 2608.11546首次发表:更新:

发表机构

Korea Institute of Science and Technology; University of Science and Technology; Korea University; Gachon University; Kyung Hee University(韩国科学技术研究院; 科技大学; 高丽大学; 嘉泉大学; 庆熙大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有艺术图像合成方法难以捕捉艺术家全局风格的问题,提出多对一的全局风格迁移范式及对应的全局风格引导、内容对齐引导方法,在WikiArt上验证了其在风格保真度等指标上的优势。

AI 中文摘要

艺术图像合成旨在重现目标艺术家富有表现力的视觉特征,但现有方法往往无法捕捉艺术家的全局风格。传统风格迁移方法以一对一的方式将一幅或少量参考艺术作品的风格迁移到内容图像,虽对作品级风格化有效,但难以表征艺术家更广泛的风格分布。以艺术家名称为条件的文本到图像扩散模型(如“梵高风格的~”)灵活性更强,但常受文本诱导偏差影响,仅复刻少数标志性作品的图案。为解决这些局限,我们提出全局风格迁移(Global Style Transfer,GST)这一艺术图像合成范式,以多对一的方式聚合目标艺术家的多幅作品,将其共享的全局风格迁移至单张内容图像。针对GST,我们提出全局风格引导(Global Style Guidance,GSG),其在固定提示词下,于扩散模型的中间特征空间(h空间)学习残差全局风格偏移;通过纯从视觉统计中学习艺术家级风格语义,GSG缓解了依赖文本的艺术偏差。我们进一步提出内容对齐引导(Content Alignment Guidance,CAG),这是一种无需训练的感知引导机制,可保留内容图像的语义结构,同时允许艺术家特定的几何变形。在WikiArt数据集上的实验表明,与现有风格迁移及基于扩散的艺术合成方法相比,GST在风格保真度、内容保留度和输出多样性方面均表现更优。

英文摘要

Artistic image synthesis aims to recreate the expressive visual identity of a target artist, yet existing methods often fail to capture an artist's global style. Conventional style transfer methods transfer the style of one or a few reference artworks to a content image in a One-to-One manner, making them effective for artwork-level stylization but limited in representing the broader stylistic distribution of an artist. Text-to-image diffusion models conditioned on artist names, such as '~ in Van Gogh style', offer greater flexibility, but they often suffer from text-induced bias and reproduce patterns from only a few iconic works. To address these limitations, we introduce Global Style Transfer (GST), an artistic image synthesis paradigm, in a Many-to-One manner, that aggregates multiple artworks from a target artist and transfers their shared global style to a single content image. For GST, we propose Global Style Guidance (GSG), which learns a residual global style offset in the intermediate feature space, or h-space, of a diffusion model under a fixed prompt. By learning artist-level style semantics purely from visual statistics, GSG mitigates text-dependent artistic bias. We further propose Content Alignment Guidance (CAG), a training-free perceptual guidance mechanism that preserves the semantic structure of the content image while allowing artist-specific geometric deformation. Experiments on WikiArt demonstrate that GST achieves superior stylistic fidelity, content preservation, and output diversity compared to existing style transfer and diffusion-based artistic synthesis methods.

CommentsPublished at ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑