AI 中文总结
本研究提出一种仅基于瓷砖补全训练的Transformer离散扩散模型,无需求解器即可实现77.4%的推箱子谜题可解率,剩余失败案例中94.5%可通过移除单个墙壁变得可解。
AI 中文摘要
判定推箱子(Sokoban)谜题是否可解是PSPACE完全问题(Culberson,1997):其解的长度可能呈指数级增长,且不存在可快速验证的简短证明。可解性也是一种脆弱属性,因为即使是单个位置错误的墙壁,也可能悄无声息地让整个谜题变得不可解。本研究表明,仅基于瓷砖补全训练、未使用求解器、奖励函数或可解性标签的Transformer型离散扩散模型,可达到77.4%的可解率,剩余失败案例中有94.5%可通过移除单个墙壁变得可解。换言之,一种依赖全局、搜索的属性可从局部训练目标中涌现:仅被训练用于填充被掩码单元格的模型,习得了其从未被训练过的可解性。自回归模型分解为$p(c_k \mid c_1 \dots c_{k-1})$,意味着固定顺序,始终以前缀为条件。掩码扩散则不同:它隐藏随机子集的单元格并学习$p(c_k \mid \text{任意子集})$,因此在生成时可按任意顺序揭示单元格,每个单元格均以棋盘上已放置的所有内容为条件。谜题的难度恰恰来自这种非局部交互:网格某一部分的决策会完全约束另一部分的可行方案。因此,不被固定顺序束缚的生成器,与该问题的结构匹配度优于自回归模型。训练流程改编自MD4(Shi等,2024),数据集为DeepMind的Boxoban(Guez等,2019)。训练好的模型及谜题生成说明已公开。
英文摘要
Deciding whether a Sokoban puzzle is solvable is PSPACE-complete (Culberson, 1997): solutions can be exponentially long and there is no short certificate to check. Solvability is also a fragile property, since even a single misplaced wall can silently render an entire puzzle unsolvable. In this work, we show that a transformer-based discrete diffusion model trained purely on tile completion, with no access to solvers, rewards, or solvability labels, achieves a solvability rate of 77.4%, with 94.5% of the remaining failures rendered solvable by removing a single wall. In other words, a global, search-heavy property follows from a local training objective: trained only to fill in masked cells, the model inherits solvability it was never trained on. An autoregressive model factorizes as $p(c_k \mid c_1 \dots c_{k-1})$, meaning a fixed order, always conditioned on a prefix. Masked diffusion does not: it hides a random subset of cells and learns $p(c_k \mid \text{any subset})$, so at generation time it can reveal cells in any order, each one conditioned on everything already placed, wherever it sits on the board. A puzzle's difficulty comes from exactly this kind of non-local interaction, a decision in one part of the grid constraining what will work somewhere else entirely. A generator that is not locked into a single fixed order is therefore a better structural match for the problem than one that is. The training pipeline is adapted from MD4 (Shi et al., 2024) and the dataset is DeepMind's Boxoban (Guez et al., 2019). The trained model and instructions for generating puzzles are publicly available.