棋局解释是否反映模型决策?LLM推理忠实度的行为与词元级测试
Do Chess Explanations Reflect Model Decisions? Behavioral and Token-Level Tests of LLM Reasoning Faithfulness
浏览论文内容
中文总结 AI 辅助
本研究通过行为与词元级测试,证明LLM在国际象棋中的流畅解释虽能影响动作偏好,但仅提供有限的推理忠实度证据,语言合理性与决策正确性相互独立。
中文摘要 AI 辅助
大型语言模型能够为国际象棋走法生成流畅的解释,但看似合理的语言并不一定反映决策背后的推理过程。我们在国际象棋中研究这一问题,因为棋盘状态完全可观测、合法动作可枚举、走法质量可独立评估。在200个Lichess残局谜题中,我们使用走法可恢复性、解码器侧控制以及合法候选走法的词元级评分来测试解释。未掩蔽的解释使生成的走法易于恢复,但在移除显式走法提示后,这一优势急剧下降。在严格掩蔽下,解释相对于仅棋盘状态仅提供微小且依赖解码器的增益。词元级评分显示,解释仍能改变走法偏好:来自其他谜题的随机但合理的解释会降低正确走法的概率,表明无关的推理文本并非被简单忽略。我们还发现,可识别的残局模式能使生成的走法更易恢复,但并不能可靠地提高走法正确性。综合来看,这些结果表明语言合理性、与生成动作的一致性以及解决方案正确性是截然不同的属性。流畅的棋局解释能影响动作偏好并支持连贯的走法叙述,同时仅提供有限的忠实推理证据。
英文摘要
Large language models can produce fluent explanations for chess moves, but plausible language does not necessarily reflect the reasoning behind a decision. We study this question in chess, where the board state is fully observable, legal actions can be enumerated, and move quality can be evaluated independently. Across 200 Lichess endgame puzzles, we test explanations using move recoverability, decoder-side controls, and token-level scoring of legal candidate moves. Unmasked explanations make generated moves easy to recover, but this advantage drops sharply after explicit move hints are removed. Under strict masking, explanations provide only small and decoder-dependent gains over the board state alone. Token-level scoring shows that explanations can nevertheless alter move preferences: random but plausible explanations from other puzzles reduce the probability of the correct move, indicating that irrelevant reasoning text is not simply ignored. We also find that recognizable endgame motifs can make generated moves easier to recover without reliably improving move correctness. Together, these results show that linguistic plausibility, consistency with a generated action, and solution correctness are distinct properties. Fluent chess explanations can influence action preferences and support a coherent move narrative while providing only limited evidence of faithful reasoning.
发表机构
- Lucerne University of Applied Sciences and Arts(卢塞恩应用科学与艺术大学)
- Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。