发表机构
University of Geneva(日内瓦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出通过直接采样干净预测或引入自我纠正训练,解决连续扩散模型在约束离散任务中早期错误难以撤销的问题,显著提升数独等任务的有效性。
AI 中文摘要
去噪扩散概率模型(DDPMs)通过从噪声开始并反复去噪来生成样本,同时保持每次更新接近当前的噪声状态。这种行为在许多连续领域中有效,但对于全局约束的离散任务(如数独、图连通性、拉丁方和N皇后)其作用尚不明确。在此类设置中,早期的离散错误可能难以撤销。因此,标准扩散采样可能会保留早期错误,即使模型的干净预测具有信息量。我们将标准采样器与直接基于模型干净预测的采样进行比较。在不重新训练的情况下,这一单一改动将数独的有效性从31%提升至95%,并在其他离散任务中取得一致改进。我们假设,保持接近当前噪声状态是有害的,因为反向轨迹可能偏离模型训练所依据的前向加噪分布。为减少这种训练-测试不匹配,我们进一步引入了自我纠正训练,该训练使模型暴露于自身的预测,从而提高对推理过程中出现错误的鲁棒性。这大幅提升了标准采样器的性能。我们的结果表明,连续扩散模型能够学习非平凡的全局约束,但离散推理任务需要更好的训练与推理对齐:要么通过减少对早期决策承诺的采样器,要么通过训练教会模型纠正自身的推理时错误。
英文摘要
Denoising Diffusion Probabilistic Models (DDPMs) generate samples by starting from noise and repeatedly denoising while keeping each update close to the current noisy state. This behavior is effective in many continuous domains, but its role is less clear for globally constrained discrete tasks, such as Sudoku, graph connectivity, Latin squares, and N-queens. In such settings, early discrete errors can be difficult to undo. As a result, standard diffusion sampling may preserve early mistakes, even when the model's clean predictions are informative. We compare standard samplers to sampling directly from the model's clean prediction. Without retraining, this single change improves Sudoku validity from 31% to 95%, with consistent gains across the other discrete tasks. We hypothesize that staying close to the current noisy state is harmful because the reverse trajectory can drift off the forward noising distribution the model was trained on. To reduce this train-test mismatch, we further introduce self-correction training, which exposes the model to its own predictions, improving robustness to errors that arise during inference. This substantially improves the performance of standard samplers. Our results suggest that continuous diffusion models can learn nontrivial global constraints, but discrete reasoning tasks require better alignment between training and inference: either through samplers that reduce commitment to early decisions, or through training that teaches the model to correct its own inference-time errors.
CommentsPreprint. The main results of this work were obtained by May 2026