AI 中文总结
本研究提出围棋死活问题基准TsuGO,发现现有LLM的推理搜索组织能力远逊于神经引导的KataGo,指出搜索组织是LLM推理评估的缺失维度。
AI 中文摘要
对大语言模型(LLM)推理的评估正从最终答案的准确性转向过程层面的评估,但现有方法仍未能捕捉到模型如何规划推理路径、分配推理资源——即如何组织搜索。现有过程层面的方法聚焦于思维链(Chain-of-Thought,CoT)的连贯性与冗余性,且大多数基准任务仅有单一目标,可通过推导、工具使用等静态能力解决,导致搜索组织这一关键维度未被测量。我们提出TsuGO,这是一个用于评估LLM推理中搜索效率的过程层面推理基准,通过围棋死活问题(tsumego)实现。这类问题提供了封闭且可验证的解空间,具有内在的对抗性结构,使得候选生成、响应检查、分支比较及回溯成为推理的必要组成部分,而非偶然的轨迹模式。通过约束解空间,TsuGO将领域知识与搜索组织解耦,将CoT解析为结构化搜索树,并报告搜索效率、令牌效率及其他诊断指标与可视化结果。实验表明,当前LLM在稳定解决围棋死活问题方面仍存在较大差距:更强的模型通过更早找到正确候选并在有效分支上持续投入精力取得成功,但大多数模型的行为更接近无引导的搜索算法,而非神经引导的KataGo。更长的CoT或更高的令牌效率并不一定意味着更好的搜索。我们的结果表明,搜索组织与推理资源分配是LLM推理评估中缺失的维度。
英文摘要
The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organize search. Prior process-level methods focus on the coherence and redundancy of chain-of-thought (CoT), and most benchmark tasks have a single objective solvable by static capabilities such as derivation and tool use, leaving search organization unmeasured. We introduce TsuGO, a process-level reasoning benchmark for evaluating Search Efficiency in LLM reasoning through Go life-and-death problems. These problems provide closed and verifiable solution spaces with an inherent adversarial structure, making candidate generation, response checking, branch comparison, and backtracking necessary parts of reasoning rather than incidental trace patterns. By constraining the solution space, TsuGO disentangles domain knowledge from search organization, parses CoT into a structured search tree, and reports Search Efficiency together with Token Efficiency and other diagnostic metrics and visualizations. Experiments show that current LLMs remain far from stable tsumego solving: stronger models succeed by finding the correct candidate earlier and sustaining effort on productive branches, but most models still behave much closer to unguided search algorithms than to neural-guided KataGo. Longer CoT or higher Token Efficiency does not necessarily imply better search. Our results identify search organization and reasoning-resource allocation as missing dimensions in LLM reasoning evaluation.
Comments23 pages, 12 figures, 20 tables, 2 algorithms