发表机构
Indraprastha Institute of Information Technology Delhi; Adobe Research(因德拉普拉斯信息技术学院德里分校; 奥多比研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有参考图像生成方法无法处理多对象复杂场景、缺乏属性级控制的问题,提出 RefDiT 框架,通过局部属性引导实现更精准的参考图像生成。
AI 中文摘要
个性化模型在少量主体参考的引导下生成新图像,而风格迁移方法旨在生成与参考图像衍生的全局风格对齐的图像。近期方法在参考图像包含单个对象时表现良好,能有效捕获涵盖所有隐式属性的全局风格。但当应用于包含多个具有不同属性特征对象的复杂真实场景时,这些方法因采用全局级引导,无法在参考图像中定位相关元素,全局引导限制了其基于参考图像局部属性生成新图像的能力。此外,现有方法通常采用单个标识符 token 捕获参考图像的所有细节,导致缺乏独立的属性级控制。受这些局限启发,我们提出 RefDiT,一种用于参考引导图像生成的新型框架。RefDiT 以参考图像、文本提示和可选的用户提供的引导上下文为输入,采用基于局部元素属性的局部区域引导。它通过对标识符 token 进行属性级分解,从参考图像构建属性感知的条件信号,并在推理提示中执行上下文调整,以训练基于扩散变换器(DiT)的生成模型的低秩适配器(LoRA)模块。RefDiT 学习标识符 token 与参考图像局部区域之间的对应关系,实现更有效的局部引导。
英文摘要
Personalization models generate new images guided by a few subject references, while style transfer methods aim to produce images aligned with a global style derived from a reference image. Recent approaches perform well when the reference image contains a single object, effectively capturing a global style that encompasses all implicit attributes. However, when applied to complex real-world scenes containing multiple objects with distinct attribute characteristics, these methods, due to their global-level guidance, fail to localize relevant elements in the reference image. The global guidance restricts their ability to generate new images based on the local attributes in the reference image. Moreover, existing methods typically employ a single identifier token to capture all details from the reference, resulting in a lack of individual, attribute-level control. Motivated by these limitations, we propose RefDiT, a novel framework for reference-guided image generation. RefDiT takes as input a reference image, a text prompt, and an optional user-provided guidance context. RefDiT employs local region guidance using the attributes of local elements. It constructs an attribute-aware conditioning signal from the reference image by performing attribute-level decomposition of the identifier token and performs context adjustment in the inference prompt to train low-rank adapter (LoRA) blocks of a diffusion transformer (DiT)-based generative model. RefDiT learns the correspondence between identifier tokens and local regions in the reference image, enabling more effective local guidance.