arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.06854cs.LGcs.AIcs.GT

关于使轻量级博弈代理强大的因素的金标准研究

A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong

  • University of Southern California(美国南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

Nima Kelidari, Mohammadsaeed Haghi, Mahdi Salmani

AI总结:

研究不完美信息纸牌游戏中使轻量级博弈代理强大的因素,构建专家代理作标准,经超百次运行确定有效因素,添加基线并应用于另一游戏,得出通用方法可训练有竞争力代理且无需专家训练。

AI中文摘要:

强化学习代理在不完美信息纸牌游戏中的强度取决于其训练对手,且难以评级。为此构建了基于规则的 Gin Rummy 专家代理作为衡量标准。通过超百次运行,确定了使轻量级代理变强的因素,如信任区域更新等有效,一些方法无效。还添加基线并应用于 Leduc Hold'em。结果是一个轻量级、通用的方法,能训练有竞争力的代理且无需专家训练。

英文摘要:

Reinforcement learning agents for imperfect-information card games are only as strong as the opponents they train against, and they are hard to grade, since they beat a random opponent over 99 percent of the time and only tie copies of themselves. So we build a strong, fixed, rule-based expert for Gin Rummy and use it only as a yardstick, never for training. It beats every agent we trained 70 to 99 percent of the time. Across more than a hundred runs, we isolate what makes a lightweight agent stronger. Trust region updates, a well-aimed reward, a curriculum of tougher opponents, warm starting, and keeping the best checkpoint all help, and stacking them lifts a self-play champion from about 30 to 36 percent against the expert. Several ideas did not pay off. Short-term and longer-term reward shaping, learned state embeddings, imitation and DAgger, and a live large language model opponent were each unhelpful, too slow, or too heavy to train at scale. Comparing MLP, convolutional, set-based, attention, and recurrent encoders shows that extra capacity does little to break the ceiling, suggesting the limit is information rather than network size. We add standard baselines (neural fictitious self-play and information set Monte Carlo search) and confirm the approach carries over to Leduc Hold'em, where the optimum is computable. The result is a lightweight, game-agnostic recipe that trains competitive agents without training on the expert, for any game a small model can handle, reported with robust statistics and released as a reusable package.

补充信息

↑