发表机构
Princeton University; Sentient Labs; University of Washington(普林斯顿大学; Sentient Labs; 华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出ACT-Eval框架,评估工具增强型LLM的象棋评论,发现事实幻觉普遍存在,工具可提升事实正确性但战略战术覆盖度仍有限。
AI 中文摘要
在国际象棋等领域,超人机游戏引擎已让专家级评估变得唾手可得,但它们仅传递事实,缺乏能让专家和非专家都受益的自然语言解释。大语言模型(LLM)原则上可填补这一空白,却因领域特定知识有限常出现幻觉,且标准的参考基准或LLM作为评判框架无法可靠检测这些错误。本研究提出ACT-Eval评估框架,将象棋评论拆解为原子主张,通过引擎支持的工具和专家标注的黄金参考评估其事实正确性、概念覆盖度及走法质量判断。我们发布包含325个局面-走法对的基准,涵盖教学、锦标赛及关键局面,其中125个局面带有专家验证的黄金原子主张,还包含五类错误分类。评估主流专有和开放权重模型后发现,象棋评论中的事实幻觉仍普遍存在:未使用工具的GPT-5.4产生错误子主张的概率为22.0%,而规模更小的开放权重模型超过40%。尽管工具增强大幅提升了事实正确性和走法质量评估,但所有模型对专家战略与战术思想的覆盖度仍有限。人类校准显示,ACT-Eval的事实判断处于人类间一致性的观测范围内,其覆盖度得分与人类对战略完整性的评估高度相关。
英文摘要
Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to experts and non-experts alike. Large language models could, in principle, bridge this gap, but they frequently hallucinate due to limited domain-specific knowledge, and standard reference-based or LLM-as-a-judge frameworks cannot reliably detect these errors. In this work, we present ACT-Eval, an evaluation framework that decomposes chess commentary into atomic claims and routes them to engine-supported tools and expert-annotated gold references to assess factual correctness, conceptual coverage, and move-quality judgment. We release a benchmark of 325 position--move pairs spanning pedagogical, tournament, and critical positions, including 125 positions with expert-verified gold atoms and a five-class error taxonomy. Evaluating leading proprietary and open-weight models, we find that factual hallucinations remain pervasive in chess commentary: GPT-5.4 without tools produces incorrect sub-claims 22.0% of the time, while smaller open-weight models exceed 40%. Although tool augmentation substantially improves factual correctness and move-quality assessment, coverage of expert strategic and tactical ideas remains limited across all models. Human calibration shows that ACT-Eval's factual judgments fall within the observed range of inter-human agreement, while its coverage scores correlate strongly with human assessments of strategic completeness.
Comments23 pages, 6 figures