发表机构
Wellcome Sanger Institute; Cambridge Stem Cell Institute, University of Cambridge; Cambridge Centre for AI in Medicine, University of Cambridge; University of Oxford; Apple; University of Helsinki; Institute for Molecular Medicine Finland (FIMM), University of Helsinki; School of Computing and Information Systems, University of Melbourne; Department of Medicine, University of Cambridge(威康桑格研究所; 剑桥大学剑桥干细胞研究所; 剑桥大学剑桥人工智能医学中心; 牛津大学; 苹果公司; 赫尔辛基大学; 赫尔辛基大学芬兰分子医学研究所; 墨尔本大学计算与信息系统学院; 剑桥大学医学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ORCA方法,通过将扩散Transformer的潜变量与视觉编码器的低秩目标对齐,并利用T5和CLIP嵌入间的残差参数化正交基,以辅助损失解决文本到图像扩散模型的组合性失败,在多个骨干上以更低训练成本提升FID和GenEval。
AI 中文摘要
文本到图像扩散模型在组合性提示上可预测地失败:属性绑定到错误的物体,空间关系颠倒,多物体场景丢失计数。最近的架构已经用T5编码器增强CLIP,正是因为CLIP的对比嵌入丢失了组合结构,然而这些失败仍然存在。我们认为绑定问题因此不是信息缺失的问题,而是信息错位的问题:文本编码器保留了组合结构,但处于一个由语言建模而非视觉塑造的表示空间中,而去噪目标并不直接奖励对齐两者。我们表明这种对应关系可以作为显式训练信号提供,相关的跨模态信息集中在自监督视觉特征的低秩子空间中,并且提供它可以作为单一辅助损失折叠到扩散训练中。我们的方法ORCA(正交残差组合对齐)将扩散Transformer的潜变量与从冻结视觉编码器导出的低秩目标对齐,通过一个预测器,其正交基由T5和CLIP嵌入之间的学习残差参数化,这为选择视觉读出子空间提供了依赖于提示的信号。我们证明了在给定秩下可恢复的跨模态信息受视觉编码器协方差在顶部组件中的谱质量的限制。在三个扩散Transformer骨干(DiT-B/2, DiT-L/2, U-ViT-L)上,ORCA在零推理时间成本下优于vanilla和REPA基线;在DiT-L/2上,它在200K步达到FID 16.65和GenEval 0.291,以一半的训练成本超过了最强的400K基线,最大的增益集中在属性绑定、空间关系和多物体提示上。
英文摘要
Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist. We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two. We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal information is concentrated in a low-rank subspace of self-supervised visual features, and that supplying it can be folded into diffusion training as a single auxiliary loss. Our method, ORCA (Orthogonal Residual Compositional Alignment), aligns the latent of a diffusion transformer with a low-rank target derived from a frozen visual encoder, through a predictor whose orthogonal basis is parameterised by a learned residual between T5 and CLIP embeddings, which provides a prompt-dependent signal for selecting the visual readout subspace. We prove that the cross-modal information recoverable at a given rank is bounded by the spectral mass of the visual encoder's covariance in the top components. Across three diffusion-transformer backbones (DiT-B/2, DiT-L/2, U-ViT-L), ORCA improves FID and GenEval over both vanilla and REPA baselines at zero inference-time cost; on DiT-L/2 it reaches FID 16.65 and GenEval 0.291 at 200K steps, exceeding the strongest 400K baseline at half the training cost, with the largest gains concentrated on attribute binding, spatial relations, and multi-object prompts.
CommentsAccepted at NeurIPS 2026. 26 pages, 4 figures