arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

会下棋并能解释其走法的语言模型

Language Models that Play Chess and Explain Their Moves

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

arXiv 2610.03695首次发表:更新:

发表机构

Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Queen,一个40亿参数的国际象棋语言模型,通过编码器-解码器架构和迭代蒸馏算法,在达到特级大师棋力的同时生成流畅解释,经七次迭代提升超900 Elo,超越前沿模型。

AI 中文摘要

现代国际象棋引擎是沉默的专家:它们以超人类水平下棋,但不会为其走法提供解释。另一方面,语言模型(LMs)能够生成听起来合理的解释,但其较弱的棋力限制了其解释的实用性。我们推出了Queen,一个具有40亿参数的国际象棋语言模型,它能在以典型特级大师水平下棋的同时解释其走法和计划。我们的新颖框架通过互补组件实现领域特定推理:一个编码器-解码器架构和一个迭代蒸馏算法。该架构通过交叉注意力将沉默的专家国际象棋编码器与经过指令调优的语言模型集成,我们通过问答课程训练该模型,以从编码器的表示中提取国际象棋概念。基于这个领域适配的模型,我们通过贝尔曼更新的自然语言类比迭代改进其解释:模型分析其顶级候选走法后的局面,并将它们整合为对当前局面的解释,然后将其蒸馏回模型中。经过七次迭代,我们的模型获得了超过900分的Elo提升(从1782到2697),在棋力和谜题准确率上大幅超越所有前沿模型,尽管其参数数量少了三个数量级。此外,基于语言模型的评估显示,我们的解释流畅,且在连贯性上接近GPT-5.6-Sol(高)。我们架构和训练过程的通用性表明,一种将语言模型应用于存在沉默专家编码器的领域(如游戏、机器人和计算机使用)的配方。

英文摘要

Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter chess-language model that can explain its moves and plans while playing at the level of a typical Grandmaster. Our novel framework enables domain-specific reasoning through complementary components: an encoder-decoder architecture and an iterative distillation algorithm. This architecture integrates a silent expert chess encoder with an instruction-tuned LM through cross-attention, which we train via a question-answering curriculum to extract chess concepts from the encoder's representations. Building on this domain-adapted model, we iteratively improve its explanations with a natural-language analog of the Bellman update: the model analyzes the positions after its top candidate moves and consolidates them into an explanation of the current position, which is then distilled back into the model. Over seven iterations, our model gains over 900 Elo points (1782 to 2697), substantially surpassing all frontier models on both playing strength and puzzle accuracy, despite containing three orders of magnitude fewer parameters. Furthermore, LM-based evaluations show that our explanations are fluent and approach GPT-5.6-Sol (high) in coherence. The generality of our architecture and training procedure suggests a recipe for applying language models to domains where silent expert encoders are available, like games, robotics, and computer use.

CommentsCode available at https://github.com/queen-project/queen

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑