发表机构
Dartmouth College(达特茅斯学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RefRoute通过紧凑残差条件化和空间路由减少多参考图像生成中的token数量与注意力开销,在ManyRef100上显著优于FLUX基线,实现高效可扩展生成。
AI 中文摘要
多参考图像生成需要在将多个主体组合成连贯场景的同时保持其外观。然而,现有的扩散Transformer通常将参考图像编码为密集的视觉token网格,并通过全局注意力联合处理,导致随着参考图像数量和分辨率的增加,条件化成本不断上升。我们提出了RefRoute框架,通过两种互补机制解决参考表示成本和注意力开销问题。紧凑残差条件化将低分辨率潜变量token与从全分辨率像素中提取的轻量级残差特征相结合,在保留细粒度外观线索的同时减少参考token数量。条件路由和注意力路由将参考token与其分配的目标区域对齐,并限制跨参考交互,同时允许超出区域边界的选择性参考访问以进行场景整合。我们进一步引入了RefRoute-Data用于训练多参考生成模型,以及ManyRef100基准,该基准涵盖人物、物体和混合组合,包含10-17个参考图像。经过多参考微调后,RefRoute在ManyRef100上实现了36.06的总体加权Ref-VIEScore,而FLUX.2-Klein-9B仅为8.88。单独的推理成本评估显示,随着参考数量的增加,延迟增长显著放缓:在16个参考图像时,我们的50步和4步配置分别比相应的FLUX基线实现了18.3倍和14.2倍的加速。这些结果确立了紧凑参考表示和空间路由注意力作为可扩展多参考图像生成的有效方法。
英文摘要
Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene. However, existing diffusion transformers commonly encode references as dense visual token grids and jointly process them with global attention, making conditioning increasingly expensive as the number and resolution of references grow. We present RefRoute, a framework that addresses both reference representation cost and attention overhead through two complementary mechanisms. Compact residual conditioning combines low-resolution latent tokens with lightweight residual features extracted from full-resolution pixels, reducing reference token counts while retaining fine-grained appearance cues. Condition routing and attention routing align reference tokens with their assigned target regions and restrict cross-reference interactions, while allowing selective reference access beyond region boundaries for scene integration. We further introduce RefRoute-Data for training many-reference generation models and ManyRef100, a benchmark spanning human, object, and mixed compositions with 10-17 references. After many-reference fine-tuning, RefRoute achieves an overall Weighted-Ref-VIEScore of 36.06 on ManyRef100, compared with 8.88 for FLUX.2-Klein-9B. Separate inference-cost evaluations show substantially slower latency growth as the reference count increases: at 16 references, our 50-step and 4-step configurations achieve $18.3\times$ and $14.2\times$ speedups over their corresponding FLUX baselines, respectively. These results establish compact reference representations and spatially routed attention as an effective approach to scalable many-reference image generation.
Comments19 pages. Wanning He and Yuyao Zhang contributed equally and share first authorship