发表机构
Institute of Applied Computer Science, Lodz University of Technology(罗兹理工大学应用计算机科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对无搜索国际象棋网络,用先验引导探索替换原熵奖励,经约2000步微调后提升了谜题与四步将死准确率,且发现战术准确率与博弈强度存在分离。
AI 中文摘要
无搜索国际象棋网络通过模仿更强的教师模型,仅需一次前向传播即可达到人类大师级水平:最强的教师模型是Leela Chess Zero(Lc0)发布的Chessformer,它蒸馏了AlphaZero风格蒙特卡洛树搜索(MCTS)的访问计数。模仿搜索对于无搜索博弈而言是糟糕的代理,因此我们通过自玩强化学习(RL)对单播强度进行微调。其探索通常由熵奖励提供,即与均匀分布的反向Kullback-Leibler(KL)散度。我们将其替换为面向网络自身MCTS先验的正向、覆盖质量的KL散度(先验引导探索),使探索覆盖先验判定为有前景的走法,并将其与由价值头结果不确定性设定的熵自适应采样温度配对,该温度在局面确定后会变得更尖锐。在约2000步内,我们在10万道谜题套件上将谜题准确率从93.9%提升至94.9%,将四步将死准确率从77%提升至81%,同时保持无搜索强度等于或略高于基线。在匹配计算范围内同时测量战术准确率和博弈强度时,我们发现二者存在分离:准确率提升处于1个百分点的区间内,而等级分则跨越基线;仅在谜题上微调的对照组取得了本研究最大的战术提升,但损失了约260 Elo;更好的谜题求解者未必是更强的棋手。分布层面的测量显示了锚定的作用:若无正则化,自玩会崩溃为单一走法线,而新求解的谜题是那些先验保留了胜势走法的差一点就能解出的谜题。正向-KL先验位居等级分榜首,与反向-KL锚定模型统计上持平,该锚定模型的集中程度是前者的两倍,但会丢失质量覆盖先验保留的最难解谜题。
英文摘要
Searchless chess networks reach human master strength from a single forward pass by imitating a stronger teacher: the strongest, Leela Chess Zero's (Lc0) released Chessformer, distills the visit counts of an AlphaZero-style Monte Carlo Tree Search (MCTS). Imitating a search is a poor proxy for playing without one, so we fine-tune for single-pass strength with self-play reinforcement learning (RL). Its exploration is usually supplied by an entropy bonus, the reverse Kullback-Leibler (KL) divergence to uniform. We replace it with a forward, mass-covering KL toward the network's own MCTS prior (prior-directed exploration), so exploration covers the moves the prior judges promising, and pair it with an entropy-adaptive sampling temperature, set by the value head's outcome uncertainty, that sharpens once a position is decided. In about two thousand steps it raises puzzle accuracy from 93.9% to 94.9% on a 100,000-puzzle suite and mate-in-four accuracy from 77% to 81% while holding searchless strength at or slightly above the base. Measuring tactical accuracy and playing strength together across a matched-compute sweep, we find the two dissociate: accuracy gains fall in a one-point band while ratings straddle the base, and a control fine-tuned on puzzles alone posts the study's largest tactical gains while shedding roughly 260 Elo; a better puzzle-solver is not thereby a stronger player. Distribution-level measurements show what anchoring buys: without a regularizer self-play collapses onto a single line of play, and the puzzles newly solved are the near misses whose winning move the prior kept alive. The forward-KL prior tops the rating ladder, statistically tied with a reverse-KL anchor that concentrates twice as hard and drops the hardest solutions the mass-covering prior keeps in support.