arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03660cs.LGcs.AI

局部更新、全局学习(LUGL):与非增量学习器博弈

Local Updates, Global Learning (LUGL): Playing Games with non-incremental Learners

David Milec, Spyridon Samothrakis, Michael Fairbank, Dennis J. N. J. Soemers

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出LUGL框架,使LightGBM等非增量学习器可用于强化学习博弈场景,在多类完美与不完美信息博弈中性能优于或相当DQN、DeepCFR,打破博弈中神经网络的主导偏见。

中文摘要 AI 辅助

神经网络(NNs)在强化学习(RL)中的主导地位部分源于其增量学习能力,该能力天然适配自博弈训练的在线非平稳特性。然而,梯度提升树(如LightGBM)被公认为监督学习中表格数据的最优方法,在准确率和效率上常优于NNs。博弈状态本质上是表格型的——离散动作、分类牌型、结构化棋盘位置——这使其成为基于树方法的理想候选。我们提出LUGL(Local Updates, Global Learning)框架,该框架将数据收集与模型拟合解耦,使梯度提升树(GBTs)等非增量学习器能在原本因分布偏移而失效的RL场景中运行。LUGL交替进行两个阶段:局部更新阶段,智能体进行自博弈对局并在有限表中积累表格型更新(Q值、V值、策略或遗憾值);全局学习阶段,利用该表训练函数逼近器以泛化到未见过的状态,之后重置该表。我们在四个标准完美信息博弈(井字棋、四子棋、奥赛罗棋、Hex棋)和五个不完美信息博弈(库恩扑克、德克萨斯扑克里德尔牌、说谎骰子、Goofspiel、Flop5德克萨斯扑克)中测试了该方法,结果显示其性能与DQN和DeepCFR相当或更优。实验表明,学界对博弈中NNs的强烈偏好可能缺乏依据,因为基于LightGBM的智能体在所有测试基准上均取得了相当或更优的性能。

英文摘要

The dominance of Neural Networks (NNs) in RL is partially due to their incremental learning capability, which naturally suits the online, non-stationary nature of self-play training. However, gradient-boosted trees like LightGBM are widely recognised as the state of the art for tabular data in supervised learning, often outperforming NNs in accuracy and efficiency. Game states are inherently tabular---discrete actions, categorical card identities, structured board positions---which makes them an ideal candidate for tree-based methods. We introduce LUGL (Local Updates, Global Learning), a framework that decouples data collection from model fitting, enabling non-incremental learners such as GBTs to operate in RL settings where they would otherwise fail due to distributional shift. LUGL alternates between a local updates phase, where the agent plays self-play games and accumulates tabular updates (Q-values, V-values, policies, or regret values) in a finite table, and a global learning phase, where the table is used to train a function approximator that generalises to unseen states before the table is reset. We test our approach in four standard perfect-information games (Tic-tac-toe, Connect-4, Othello, and Hex) and five imperfect-information games (Kuhn's poker, Leduc Hold'em, Liar's Dice, Goofspiel, and Flop5 Hold'em), and show that our results are competitive with or superior to DQN and DeepCFR. Our experiments demonstrate that the community's strong bias towards NNs in game-playing may be unwarranted, since LightGBM-based agents achieve competitive or superior performance across all tested benchmarks.

发表机构

  • Czech Technical University in Prague(布拉格捷克技术大学)
  • University of Essex(埃塞克斯大学)
  • Maastricht University(马斯特里赫特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑