arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02425cs.CL

找到走法并非赢得对局:XiangqiBench 用于LLM代理的闭环评估

Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

Yekun Chai, Qiwei Peng, Haoyi Xiong

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM代理在中国象棋中的闭环表现,提出XiangqiBench基准,通过119个残局测试发现转换、一致性和模拟三大差距,强调应评估闭环结果而非仅走法正确性。

中文摘要 AI 辅助

静态评估会因语言模型说出正确走法而给予其评价,但代理必须在对手回应的同时,将计划执行到经过验证的结果。我们引入了XiangqiBench,一个可执行的基准测试,用于衡量中国象棋中的这种差异:从119个由引擎或仅检查搜索支持的强制将杀战术残局开始,LLM代理必须对引擎防守方实现将杀。一个交互式REPL界面将真实走法、状态查询和前瞻模拟分开,我们在两种观察协议下记录了来自12个前沿LLM的8,568个多轮轨迹。三种看似能力的信号都高估了闭环成功。(i) 转换差距:模型在26.1%的有视觉试验中走出存储的参考首步,但这些试验中只有13.9%以胜利结束。(ii) 一致性差距:领先模型达到38.7%的pass@3,但只有5.9%的pass^3,在其曾经获胜的46个局面中,仅在7个局面中赢得全部三次试验。(iii) 模拟差距:32.3%的被接受的模拟调用因非法走法而停止,在49.3%的可比案例中,真实防守方的回应与代理模拟的线路不同;自生成的推演检查合法性,但无法预判对手。找到走法并非赢得对局:代理评估应评分闭环结果,并同时报告可靠性与覆盖率。

英文摘要

Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this difference in Chinese chess: starting from 119 tactical endgames with forced mates supported by engine or checks-only search, an LLM agent must deliver checkmate against an engine defender. An interactive REPL interface separates real moves, state queries, and forward simulation, and we record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. Three signals that look like competence each overstate closed-loop success. (i) The Conversion Gap: models play the stored reference first move in 26.1\% of Sighted trials, yet only 13.9\% of these trials end in a win. (ii) The Consistency Gap: the leading model reaches 38.7\% pass@3 but only 5.9\% pass^3, winning all three trials on 7 of the 46 positions it ever wins. (iii) The Simulation Gap: 32.3\% of accepted simulation calls stop on an illegal move, and in 49.3\% of comparable cases the real defender replies differently from the line the agent simulated; self-authored rollouts check legality but cannot anticipate the opponent. Finding the move is not winning the game: agent evaluations should score closed-loop outcomes and report reliability alongside coverage.

↑