arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26946cs.AI

围棋中结合不确定性门控的信念引导决策方法

Belief-Guided Decision Making with Uncertainty Gating in the Game of Go

Mehrad Yaghoubi, Azam Bastanfard, Abbas Jalilvand, Ashkan Rezaei

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对计算机围棋消费级硬件推理瓶颈与模型幻觉问题,提出信念引导架构,解耦策略头与信念头,结合记忆机制与门控,在有限硬件上实现职业级对弈

中文摘要 AI 辅助

近年来,AlphaZero和MuZero推动的计算机围棋领域进展高度依赖蒙特卡洛树搜索(MCTS)来修正神经网络策略的误差。尽管在大规模计算集群上效果显著,但这种依赖在消费级硬件上形成了关键瓶颈,树管理的计算成本严重限制了推理速度。此外,若不进行深度搜索,这些模型会出现幻觉,提出看似高置信度却具有战略致命性的落子。本文提出一种新颖的信念引导架构,将策略头与独立的信念头解耦。与传统价值函数不同,信念头充当内部模拟器和独立评判者,建模认知不确定性与战略稳定性。通过集成Transformer/GRU等记忆机制处理长期依赖、加入打劫规则(Ko rule),并利用门控机制过滤过度自信的策略误差,我们的模型将智能负担从运行时搜索转移到参数化的“直觉”上。实验结果表明,该方法显著提升了无搜索胜率并减少了幻觉,在无法使用大规模MCTS的有限硬件上实现了职业级对弈。

英文摘要

Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural network policy. While effective on massive computational clusters, this dependence creates a critical bottleneck on consumer-grade hardware, where the computational cost of tree management severely limits inference rates. Furthermore, without deep search, these models suffer from hallucination, proposing moves with high confidence that are strategically fatal. This paper introduces a novel Belief-Guided architecture that disentangles the Policy head from a distinct Belief head. Unlike traditional value functions, the Belief head acts as an internal simulator and independent critic, modeling epistemic uncertainty and strategic stability. By integrating memory mechanisms (Transformer/GRU) to handle long-term dependencies and the Ko rule, and utilizing a gating mechanism to filter overconfident policy errors, our model shifts the burden of intelligence from runtime search to parametric "intuition." Experimental results demonstrate that this approach significantly improves search-free win rates and reduces hallucination, enabling professional-level play on limited hardware where massive MCTS is infeasible.

补充信息

↑