AI 中文总结
本研究在受限算力下通过搜索自对弈训练国际象棋模型,经2.5天训练达3251 Elo,验证了集成系统效率工程的整体贡献。
AI 中文摘要
在有限的训练算力下,当整个学习循环为效率而精心设计时,一个AlphaZero风格的国际象棋系统能达到多强的水平?我们在单个八GPU节点上,从随机初始化开始,通过搜索自对弈训练了2.5天。由此产生的632万参数模型,在每步棋10万次搜索(估计思考时间不到五秒)的情况下,达到了3,251基准Elo [3,206, 3,297],对抗固定节点的Stockfish 13阶梯。该运行处理了325万局完整对局,涉及估计1000亿次搜索模拟,并进行了8.366亿次训练展示。我们研究了搜索分配、回放与重启状态选择、策略表示、渐进式模型规模、量化推理和吞吐量工程。在保留设计的同时,我们记录了未能改善完整学习循环或未能证明其成本合理的可行替代方案。所报告的强度是集成系统的结果,而非任何单一选择可归因的孤立Elo增益。
英文摘要
How strong can an AlphaZero-style chess system become under limited training compute when its entire learning loop is engineered for efficiency? We train from random initialization through searched self-play on a single eight-GPU node for 2.5 days. The resulting 6.32-million-parameter model reaches 3,251 benchmark Elo [3,206, 3,297] at 100,000 searches per move (estimated at under five seconds of thinking time) against a fixed-node Stockfish 13 ladder. The run ingests 3.25 million completed games, involves an estimated 100 billion search simulations, and makes 836.6 million training presentations. We investigate search allocation, replay and restart-state selection, policy representation, progressive model sizing, quantized inference, and throughput engineering. Alongside the retained design, we document plausible alternatives that failed to improve the complete learning loop or did not justify their cost. The reported strength is a result of the integrated system, not an isolated Elo gain attributable to any single choice.
Comments26 pages, 17 figures, including appendices. Code and experimental evidence: https://github.com/BertilBraun/Advanced-Techniques-in-Chess-Engines ; model artifacts: https://huggingface.co/BertilBraun/alphazero-chess ; interactive demo: https://chess.bertil-braun.de/