发表机构
Institute of Machine Learning and Neural Computation; Graz University of Technology(机器学习与神经计算研究所; 格拉茨技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散模型难以满足结构化空间推理约束的问题,提出利用无监督对象发现与关系抽象增强模型推理能力,并构建四个谜题数据集,验证了该方法在复杂推理与分布外泛化上的显著提升。
AI 中文摘要
扩散模型在图像合成方面表现出色,但在可靠满足结构化空间推理约束方面仍然存在局限。在具有隐式逻辑结构的条件数据分布建模任务中,例如由可见线索与一致解配对定义的谜题,最先进的生成模型往往倾向于逼近像素空间分布,而不学习推理所需的基础逻辑规则。为解决这一局限,我们提出了一种用于扩散模型空间推理的新框架,该框架利用无监督对象发现和对象关系抽象。我们表明,从以对象为中心的表示中提取的关系知识为扩散模型丰富了结构基元,使其能够在训练和推理过程中有效引导生成表示空间,并实现满足推理约束的条件图像生成。此外,我们引入了一个大规模生成式空间推理基准,包含四个受人类可解谜题启发的数据集。我们的结果表明,关系抽象显著提升了扩散模型在多种复杂推理任务上的推理能力,同时实现了在分布外设置中的稳健泛化。
英文摘要
Diffusion models excel at image synthesis, but they remain limited in their ability to reliably satisfy structured spatial reasoning constraints. In conditional data distribution modeling tasks with implicit logical structure, such as puzzles defined by visible clues paired with consistent solutions, state-of-the-art generative models tend to approximate pixel-space distributions without learning the underlying logical rules required for inference. To address this limitation, we present a novel framework for spatial reasoning with diffusion models that leverages unsupervised object discovery and abstractions of object relations. We show that the relational knowledge derived from object-centric representations enriches diffusion models with structural primitives, allowing them to effectively guide the generative representation space during both training and inference, and enabling conditional image generation that satisfies reasoning constraints. Additionally, we introduce a large-scale generative spatial reasoning benchmark with four datasets inspired by human-solvable puzzles. Our results show that relational abstractions significantly improve reasoning capabilities of diffusion models on a variety of complex reasoning tasks, while enabling robust generalization in out-of-distribution settings.