全景场景程序扩散Transformer
Panoptic Scene Program Diffusion Transformer
浏览论文内容
中文总结 AI 辅助
提出PSP-DiT,一种将全景场景程序作为潜在变量联合去噪的扩散Transformer,在多个基准上显著提升组合提示生成能力,尤其改善计数、属性绑定和角色敏感关系。
中文摘要 AI 辅助
现代文本到图像模型能够生成高保真图像,但在处理需要实例身份、属性归属、计数、空间排序和角色敏感关系的组合提示时仍存在困难。我们引入了全景场景程序扩散Transformer(PSP-DiT),这是一种扩散Transformer架构,将全景场景程序视为一等潜在变量,而非外部控制信号或事后解析。PSP-DiT通过耦合的Transformer流联合去噪图像潜在变量和场景程序潜在变量,同时全景锚定和循环一致性目标将对象实例、属性、关系和计数与生成图像中的视觉支持绑定。在匹配的训练和推理设置下,PSP-DiT在GenEval 2、SANEval-Simple、PSG-Score和DetailMaster上优于强大的平面文本基线,在计数、属性绑定、角色敏感关系和长结构化提示上提升最大。该方法保持了图像质量,增加了适度的推理开销,并且对不完美的场景程序保持鲁棒性。
英文摘要
Modern text-to-image models produce high-fidelity images but still struggle with compositional prompts that require instance identity, attribute ownership, counting, spatial ordering, and role-sensitive relations. We introduce Panoptic Scene Program Diffusion Transformer (PSP-DiT), a diffusion-transformer architecture that treats a panoptic scene program as a first-class latent variable rather than an external control signal or post-hoc parse. PSP-DiT jointly denoises image latents and scene-program latents through coupled transformer streams, while panoptic grounding and cycle-consistency objectives tie object instances, attributes, relations, and counts to visual support in the generated image. Under matched training and inference settings, PSP-DiT improves over a strong flat-text baseline across GenEval 2, SANEval-Simple, PSG-Score, and DetailMaster, with the largest gains on counting, attribute binding, role-sensitive relations, and long structured prompts. The method preserves image quality, adds modest inference overhead, and remains robust to imperfect scene programs.