arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用离线强化学习发现高质量国际象棋谜题

Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning

Allen Nie, Anirudhan Badrinath, Nicholas Tomlin, Timothy Dai, Carissa Yip, Rose E Wang, Emma Brunskill, Chris Piech

arXiv 2608.14851首次发表:更新:

AI 中文总结

本研究利用15亿条国际象棋谜题解答历史数据,通过离线强化学习学习谜题教学价值,能为100-1000 Elo的初学者推荐高质量谜题,尤其助力学习停滞的新手。

AI 中文摘要

学习与技能掌握需要大量刻意练习,在许多学习场景中,制作高质量教学材料往往需要高水平的领域专业知识且耗时极长。教学材料常需训练学生形成不同的思维模式,在国际象棋这类领域,谜题被用于帮助学生练习计算后续走法、识别棋盘上的已知模式。为学生提供一套谜题练习以培养不同思维模式颇具挑战,因为教师需仔细平衡不同主题以及学生需要执行的前瞻步数。像chess.com和Lichess这类热门在线平台拥有数百万级谜题,与人类专家制作的国际象棋战术谜题(初学者可从中学习宝贵见解)不同,这些谜题多为自动生成,常被认为教学价值较低。这些平台还依赖启发式方法为用户推荐练习谜题。我们利用全年用户历史数据(共15亿条谜题解答历史),借助离线强化学习的见解学习谜题的教学价值,并自动选择一组谜题以更好支持国际象棋学习者。我们通过离线策略评估表明,训练出的策略对谜题解答Elo等级分在100-1000区间的初学者有显著影响,尤其对学习增长停滞的初学者群体效果明显。我们还通过收集国际象棋专家的标注评分,对模型发现的谜题进行了定性分析。我们这套流程的成功,为未来基于通用用户交互数据理解练习项目的教学价值展现了广阔前景。

英文摘要

Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing high-quality pedagogical materials can require a high level of domain expertise and be very time-consuming. Pedagogical materials often need to train students to engage in different thinking patterns. In some domains, such as chess, puzzles are used to help students practice their skills in calculating the next moves and recognizing known patterns on a board. Giving students a practice set of puzzles to help them learn different modes of thinking is challenging because the teacher needs to carefully balance between different motifs and how many look-ahead steps a student needs to perform. Popular online platforms like Chess.com and Lichess offer players millions of puzzles. Unlike chess tactics puzzles procured by human experts, where chess beginners can learn valuable insights, these puzzles are automatically generated and often regarded as having low pedagogical value. These platforms also rely on a heuristic to recommend puzzles to users for practice. Using the user history data over an entire year, a total of 1.5 billion puzzle-solving histories, we learn the pedagogical value of a puzzle and how to automatically choose a set of puzzles to better support chess learners using insights from offline reinforcement learning. We show that using offline policy evaluation, our trained policy has significant impact on beginners with puzzle-solving Elo range of 100--1000, particularly for the group of beginners whose learning growth was stagnant. We also performed a qualitative analysis of the puzzles discovered by our model by collecting annotation ratings from expert chess players. The success of our pipeline shows promise for a future where we can understand the pedagogical values of practice items given general user interaction data.

CommentsPublished at RLC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑