arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38577cs.AIcs.LG

基于扩散模型的创意国际象棋谜题条件生成

Conditional Generation of Creative Chess Puzzles with Diffusion Models

Aatu Selkee, Severi Rissanen, Xidong Feng, Tom Zahavy, Eric Malmi

首次发表
浏览论文内容

中文总结 AI 辅助

针对语言模型在受约束创造性任务上的不足,本文提出用掩码扩散模型条件生成国际象棋谜题,通过辅助最佳走法预测和强化学习优化,显著提升唯一性与主题匹配度,并开源首个模型。

中文摘要 AI 辅助

尽管现代语言模型展现出令人印象深刻的生成能力,但它们在处理受约束的、反直觉的创造性任务时往往力不从心。为了解决这一局限性,我们将国际象棋谜题生成作为计算创造力和推理的严格测试平台,在这个领域中,改变一个棋子就可能使整个解决方案失效。我们提出了一种新颖的方法,利用掩码扩散模型进行创意国际象棋谜题的条件生成。与以往方法不同,我们的非定向扩散方法允许针对特定战术主题和部分棋盘位置进行条件设置。我们引入了一项新颖的辅助任务——同步最佳走法预测,该任务将解决方案唯一性提高了11.6%,主题条件准确率提高了2.5%。为了进一步优化解决方案唯一性和主题条件设置,我们建立了一个改编自去噪扩散策略优化(DDPO)的强化学习框架。这种强化学习训练使唯一且主题匹配的位置产出增加了89.1%。最后,我们发布了首个用于国际象棋谜题生成的开源权重模型(附录B),为可控的、创造性的生成提供了一条新途径。

英文摘要

While modern language models demonstrate impressive generative capabilities, they often struggle with constrained, counter-intuitive creative tasks. To address this limitation, we explore chess puzzle generation as a rigorous testbed for computational creativity and reasoning, a domain where altering a single piece can invalidate an entire solution. We propose a novel approach for conditional generation of creative chess puzzles using masked diffusion models. Unlike previous methods, our non-directional diffusion approach allows for conditioning on specific tactical themes and partial board positions. We introduce a novel auxiliary task of simultaneous best-move prediction, which improves solution uniqueness by 11.6% and theme-conditioning accuracy by 2.5%. To further optimize solution uniqueness and theme conditioning, we establish a reinforcement learning framework adapted from Denoising Diffusion Policy Optimization (DDPO). This RL training increases the yield of unique and theme-matching positions by 89.1%. Finally, we release the first open-weights models (Appendix B) for chess puzzle generation, offering a new pathway for controllable, creative generation.

发表机构

  • Aalto University(阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

↑