RefCompose:基于LoRA条件扩散的多参考图像生成
RefCompose: Multi-Reference Image Generation via LoRA-Conditioned Diffusion
浏览论文内容
中文总结 AI 辅助
RefCompose通过固定分辨率参考画布和双流LoRA适配器解耦布局与外观,实现恒定内存的多参考图像生成,在Dense Layout协议上优于现有方法。
中文摘要 AI 辅助
电影制作人和视觉艺术家经常需要将多个参考元素(演员、地点、道具、文化元素)组合成一个连贯的镜头,但现有工具要么无法扩展到超过少数几个参考,要么在此过程中破坏细粒度的主体身份,因为每个参考的标记化使得内存随参考数量N线性增长,且生成的内容往往偏离给定的参考而非再现它们。我们提出RefCompose,一种像素空间组合条件框架,通过一个固定分辨率的参考画布将“位置”(where)与“外观”(what)解耦,无论参考数量多少,条件大小保持不变。空间布局在推理时由冻结的扩散变换器诱导,并通过Grounding DINO提取,无需LLM或专用布局模型,而双流LoRA适配器通过独立的低秩流注入布局导出的深度图和编码画布,将几何支架与局部外观分离。在Dense Layout协议上,RefCompose在颜色、纹理、形状、空间准确性以及更高参考数量下的身份/内容保留方面,始终优于基于布局的和最先进的多参考基线,且推理内存恒定,使其成为生产规模多主体电影合成的实用构建模块。
英文摘要
Filmmakers and visual artists routinely need to compose multiple references, actors, locations, props, cultural elements, into a single coherent shot, but existing tools either fail to scale past a handful of references or destroy fine grained subject identity in the process, since per reference tokenization scales memory linearly with reference count $N$ and generated content often departs from the given references rather than reproducing them. We propose \textbf{RefCompose}, a pixel space compositional conditioning framework that decouples \emph{where} things go from \emph{what} they look like, via a single fixed resolution reference canvas that keeps conditioning size constant regardless of reference count. Spatial layout is induced at inference time from a frozen diffusion transformer and extracted via Grounding DINO, requiring no LLM or dedicated layout model, while dual stream LoRA adapters inject a layout derived depth map and the encoded canvas through separate low rank streams, disentangling geometric scaffolding from localized appearance. On the Dense Layout protocol, RefCompose consistently outperforms layout based and state of the art multi reference baselines on color, texture, shape, spatial accuracy, and identity/content preservation at higher reference counts, all with constant inference memory, making it a practical building block for multi subject cinematic composition at production scale.