arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

外观指针——扩散变压器的多模态区域控制

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch, Srinath Sridhar, Matheus Gadelha

arXiv 2607.19344首次发表:更新:

发表机构

Brown University; Adobe Research(布朗大学; Adobe 研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对可控图像生成难题,提出外观指针方法,通过区域对应网络和空间聚合机制生成并细化指针,为扩散变压器引入模态无关的局部多模态控制接口,单一模型性能达或超现有技术。

AI 中文摘要

对于创意专业人士来说,可控图像生成仍然具有挑战性,他们通常需要对材料、物体身份和空间布局进行精确的区域控制,而仅通过文本提示无法可靠地实现这一点。扩散变压器(DiTs)可以原生地摄取来自文本和图像的异构令牌,但它们缺乏确定这些令牌应在何处以及如何影响输出的机制。我们引入了外观指针,通过将文本或图像输入与用户指定的掩码对齐,在正确的空间位置将DiTs引向正确的外观线索。外观指针由区域对应网络生成,并通过空间聚合机制进行细化,使模型能够处理多个区域描述而不会显著增加令牌负载。我们的方法在不从头重新训练基础模型的情况下,为DiT中的局部多模态控制引入了第一个模态无关接口。在一系列指标上,我们的单一模型达到或超过了特定模态的现有技术方法的性能,为生成图像合成中的精确、区域感知、多模态指导提供了一条简单且可扩展的途径。

英文摘要

Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output. We introduce appearance pointers, compact tokens that guide DiTs toward the correct appearance cues at the correct spatial locations by aligning text or image inputs with user-specified masks. Appearance pointers are produced by a region correspondence network and refined through a spatial aggregation mechanism, enabling the model to handle multiple regional descriptions without significantly increasing token load. Our approach introduces the first modality-agnostic interface for localized multimodal control in a DiT without retraining the base model from scratch. Across a range of metrics, our single model reaches or surpasses the performance of modality-specific state of the art methods, offering a simple and extensible path toward precise, region-aware, multimodal guidance in generative image synthesis.

Comments38 Pages, Preprint with supplement

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑