arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoBind:用于无训练文本到图像生成的阶段感知组合绑定

CoBind: Stage-Aware Compositional Binding for Training-Free Text-to-Image Generation

Kaijie Chen, Ethan Caldwell, Mira Vossen, Julian Hartwell, Serena Whitlock, Adrian Bellamy

arXiv 2607.16307首次发表:更新:

发表机构

Mind Lab(思维实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对扩散文本到图像模型处理复杂提示失败的问题,提出CoBind框架,先解析提示为组合图,利用多种约束建立布局、绑定属性,去噪时调整引导强度,无需重新训练或注释,实验证明其在多方面有改进且视觉质量有竞争力。

AI 中文摘要

基于扩散的文本到图像模型在处理涉及多个实体、属性和关系的复杂提示时经常失败,会出现对象遗漏、属性分配错误或空间布局颠倒等问题。现有无训练方法主要强化令牌级注意力,但未明确建模哪些属性属于哪些实体,或在去噪过程中何时应实施不同约束。我们引入了CoBind,一个用于阶段感知组合绑定的无训练框架。CoBind将提示解析为实体、属性和关系的组合图。它首先利用实体完整性和关系约束建立全局布局,然后通过对比跨实体优化将属性绑定到其目标实体。在后期去噪步骤中逐渐放宽结构引导以保留纹理和视觉细节。CoBind还根据每个约束的当前满足情况调整引导强度,减少不必要的潜在更新。CoBind无需重新训练或额外注释。在T2I-CompBench++、GenEval和多个扩散主干上的实验表明,在属性绑定、空间关系和复杂组合生成方面有一致改进,同时保持有竞争力的视觉质量。

英文摘要

Diffusion-based text-to-image models often fail on complex prompts involving multiple entities, attributes, and relations, producing object omissions, incorrect attribute assignments, or reversed spatial layouts. Existing training-free methods mainly strengthen token-level attention, but do not explicitly model which attributes belong to which entities or when different constraints should be enforced during denoising. We introduce \textbf{CoBind}, a training-free framework for stage-aware compositional binding. CoBind parses a prompt into a composition graph of entities, attributes, and relations. It first establishes the global layout using entity-completeness and relation constraints, then binds attributes to their target entities through contrastive cross-entity optimization. Structural guidance is gradually relaxed in later denoising steps to preserve textures and visual details. CoBind also adapts the guidance strength according to the current satisfaction of each constraint, reducing unnecessary latent updates. CoBind requires no retraining or additional annotations. Experiments on T2I-CompBench++, GenEval, and multiple diffusion backbones show consistent improvements in attribute binding, spatial relations, and complex compositional generation while maintaining competitive visual quality.

Comments22 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑