发表机构
Warsaw University of Technology; IDEAS Research Institute; University of Amsterdam(华沙理工大学; IDEAS研究所; 阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Painter-Thinker模型,通过递归潜在推理在像素空间解决视觉推理任务,无需符号监督,显著提升数独等难题的解决率。
AI 中文摘要
扩散模型能够生成逼真的图像,但在视觉推理任务上常常失败,例如填充数独或绘制迷宫路径。当离散符号表示可用时,递归方法如微型递归模型(TRM)甚至能解决这些难题的困难实例。我们探究如何将这种推理迁移到像素空间,其中没有符号表示可用。我们提出Painter-Thinker(PaTh):一个小型递归网络(Thinker)在编码噪声图像和条件的网格化学习令牌上进行推理,在每个去噪步骤内细化潜在状态,并通过ControlNet适配器引导冻结的扩散模型(Painter)。Thinker仅使用标准重建损失进行训练,无需符号目标、求解器或验证器。PaTh解决了92.5%的困难MNIST数独谜题(先前最佳为75%)和71.2%的极端谜题(先前最佳为4.1%),参数量为1000万,而标准扩散模型为8200万。它还在迷宫、皇后问题和具有指定空间关系的CLEVR场景上有所改进,且其优势随问题规模增大而增强。诊断实验表明,PaTh能从扩散模型无法修复的注入错误中恢复,尤其是当许多单元格错误时。这些结果表明,为符号数据开发的推理机制可以无需符号监督而集成到像素空间扩散中,为在日益复杂的约束下生成数据开辟了道路。
英文摘要
Diffusion models generate realistic images but often fail on visual reasoning tasks, such as filling in a Sudoku or drawing the path through a maze. When a discrete symbolic representation is available, recursive methods such as the Tiny Recursive Model (TRM) solve even hard instances of these puzzles. We ask how such reasoning can be carried over to pixels, where no symbolic representation is available. We propose Painter-Thinker (PaTh): a small recursive network (the Thinker) reasons over a grid of learned tokens that encode the noisy image and the conditioning, refines a latent state within every denoising step, and steers a frozen diffusion model (the Painter) through ControlNet adapters. The Thinker is trained with the standard reconstruction loss alone, without symbolic targets, a solver, or a verifier. PaTh solves 92.5% of hard MNIST Sudoku puzzles (prior best 75%) and 71.2% of extreme ones (prior best 4.1%), with 10M parameters against 82M for a standard diffusion model. It also improves on mazes, Queens, and CLEVR scenes with specified spatial relations, and its advantage grows with problem size. Diagnostic experiments show that PaTh recovers from injected mistakes that the diffusion model cannot repair, especially when many cells are wrong. Together, these results show that reasoning mechanisms developed for symbolic data can be integrated into pixel-space diffusion without symbolic supervision, opening a path toward generating data under increasingly complex constraints.