arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ControlRef:基于锚定4D-RoPE的高效布局引导多实例生成

ControlRef: Efficient Layout-Guided Multi-Instance Generation via Anchored 4D-RoPE

Yunkai Yang, Yudong Zhang, Xinying Chen, Haoyuan Liang, Yizhuo Niu, Jinshuai Cheng, Kunquan Zhang, Liziyue Fang, Weitao Wan, Runmin Dong

arXiv 2608.06878首次发表:更新:

AI 中文总结

ControlRef是一种高效布局引导多实例合成框架,通过UILC注意力掩码与锚定4D-RoPE提升空间对齐,在保证视觉保真度和定位精度的同时大幅降低推理延迟与内存开销。

AI 中文摘要

布局引导的多实例生成对于多模态扩散Transformer(MM-DiTs)中的可控图像合成至关重要,但将该能力集成到统一架构中仍具挑战性。现有框架依赖冗余的全分辨率画布填充和移位旋转位置编码(Shifted-RoPE)来管理多幅参考图像,该机制会大幅增加稀疏布局下的计算开销,并破坏关键的低频旋转位置编码(RoPE)特征,造成严重的空间-频率折衷问题,模糊绝对空间对应关系。为克服这些局限,我们提出ControlRef,一个高效且精确的多实例合成框架。ControlRef采用统一实例-布局控制(UILC)注意力掩码,严格解耦实例间的语义交互并实现精确的区域绑定;为进一步促进区域级空间对齐,我们引入锚定4D-RoPE,一种新型位置编码机制,直接将token锚定到其绝对几何中心。通过将参考图像预对齐到对应边界框分辨率,将布局与参考token物理锚定到绝对几何中心,并沿z轴堆叠参考图像,锚定4D-RoPE可原生保留空间先验,且无损耗移位地缓解空间-频率折衷问题。大量实验表明,ControlRef实现了最先进的视觉保真度与定位精度,同时在稀疏布局下将推理延迟降低80%以上,在密集场景中减少50%的内存开销。

英文摘要

Layout-guided multi-instance generation is essential for controllable image synthesis in Multi-Modal Diffusion Transformers (MM-DiTs). However, integrating this capability into unified architectures remains challenging. Prior frameworks rely on redundant full-resolution canvas padding and Shifted-RoPE to manage multiple reference images. This mechanism drastically inflates computational overhead for sparse layouts and disrupts critical low-frequency RoPE features, creating a severe spatial-frequency compromise that blurs absolute spatial correspondence. To overcome these limitations, we propose ControlRef, a highly efficient and precise multi-instance synthesis framework. ControlRef utilizes a Unified Instance-Layout Control (UILC) attention mask to strictly decouple inter-instance semantic interactions and enforce precise regional binding. To further promote region-level spatial alignment, we introduce Anchored 4D-RoPE, a novel positional encoding mechanism that directly anchors tokens to their absolute geometric centers. By pre-aligning reference images to their corresponding bounding box resolutions, physically anchoring both layout and reference tokens to their absolute geometric centers, and stacking the references along the z-axis, Anchored 4D-RoPE natively preserves spatial priors and mitigates the spatial-frequency compromise without lossy shifting. Extensive experiments demonstrate that ControlRef achieves state-of-the-art visual fidelity and localization accuracy, while concurrently slashing inference latency by over 80% in sparse layouts and reducing memory overhead by 50% in dense scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑