arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

救赎自在其中:在步进蒸馏扩散模型中引发固有风格迁移

Salvation Lies Within: Eliciting Inherent Style Transfer in Step-Distilled Diffusion Models

Shengyin Sun, Yiming Li, Yingzhao Lian, Xing Li, Xingzhi Zhou, Anxin Tian, Zhili Wang, Haoyang Li, Ziqiang Cui, Chen Ma

arXiv 2610.05066首次发表:更新:

发表机构

Huawei Technologies; Hong Kong Polytechnic University; City University of Hong Kong(华为技术有限公司; 香港理工大学; 香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出StyleForge,一种无需训练的框架,通过将参考风格编译为可复用文本指令,使冻结的步进蒸馏T2I模型实现风格迁移,在保留少步生成的同时,质量得分相对提升高达29.47%。

AI 中文摘要

通过后训练来适配步进蒸馏的文本到图像(T2I)模型会带来额外的计算成本,并影响其原生的少步生成行为。这促使我们探索一条超越特定风格适配的补充路径:利用步进蒸馏T2I模型中已编码的视觉知识,通过语言来引发风格化能力。沿着这一方向,需要文本引导能够捕捉视觉属性如何共同定义一种风格,并在描绘内容变化时保持适用。为探索这一方法,我们引入了StyleForge,一个全自动、无需训练的框架,将参考风格表达为可复用的渲染指令。通过整合整体渲染特征与局部颜色和光照行为,StyleForge将参考图像中的视觉证据组织成关于目标风格应如何表达的连贯规范。该规范随后被编译成文本引导,可跨内容提示复用,使冻结的步进蒸馏T2I模型能够以参考风格渲染不同主体和场景,同时保留原生的少步生成能力。大量实验表明,与最强基线相比,生成质量得分相对提升高达29.47%,而帕累托分析表明,改进的风格化伴随着对请求内容的强遵循。

英文摘要

Adapting step-distilled text-to-image (T2I) models through post-training incurs additional computational costs and affects native few-step generation behavior. This motivates a complementary route beyond style-specific adaptation: drawing on the visual knowledge already encoded in step-distilled T2I models to elicit stylistic capabilities through language. Pursuing this direction requires textual guidance that captures how visual attributes jointly define a style and remain applicable as the depicted content changes. To explore this approach, we introduce StyleForge, a fully automatic, training-free framework that expresses reference styles as reusable rendering instructions. By integrating overall rendering characteristics with local color and lighting behavior, StyleForge organizes visual evidence from reference images into a coherent specification of how the target style should be expressed. The specification is then compiled into textual guidance that can be reused across content prompts, enabling frozen step-distilled T2I models to render different subjects and scenes in the reference style while retaining native few-step generation. Extensive experiments show relative gains of up to 29.47\% in generation quality scores over the strongest baseline, while Pareto analysis indicates that improved stylization is accompanied by strong adherence to the requested content.

Comments29 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑