实时策略游戏中结合老虎机策略选择的约束指令条件强化学习
Constrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy Games
浏览论文内容
中文总结 AI 辅助
该研究针对实时策略游戏,提出结合汤普森采样老虎机策略选择的约束指令条件PPO强化学习系统,在MicroRTS环境中,相比扁平PPO基线,对多数对手胜率显著更高。
中文摘要 AI 辅助
深度强化学习智能体在实时策略游戏中表现出色,但对训练分布之外的对手较为脆弱。将战略指令选择与已学习的单位控制分离,可针对不同对手选择不同策略,同时复用相同执行策略。这需要一个能遵循不同指令的执行器,以及用于评估其是否做到这一点的可测量标准。我们为实时策略环境MicroRTS引入了一种约束指令条件近端策略优化(PPO)策略,即执行器。离散指令在多个环境步骤中指定经济、部队编成、军事态势和工人策略的战略目标与行为要求;执行器确定用于实现这些目标的单位级动作。汤普森采样老虎机充当战略家,根据游戏内观察构建的对手策略估计值选择指令元组,而非基于对手身份。在与采用相同架构、预算、课程和自博弈联盟训练的扁平PPO基线的受控对比中,该战略家-执行器系统在训练地图上对四个最强对手中的三个胜率显著更高,包括两个最强的保留对手(胜率区间为0.55至0.97和0.01至0.34),对其余对手则无显著差异。
英文摘要
Deep reinforcement learning agents reach strong performance in real-time strategy games but can be brittle against opponents outside their training distribution. Separating strategic command selection from learned unit control allows different strategies to be selected for different opponents while reusing the same execution policy. This requires an executor that can follow different commands and measurable criteria for assessing whether it does so. We introduce a constrained command-conditioned Proximal Policy Optimization (PPO) policy, the executor, for MicroRTS, a real-time strategy environment. Discrete commands specify strategic objectives and behavioral requirements for economy, army composition, military posture, and worker policy over multiple environment steps; the executor determines the unit-level actions used to fulfill them. A Thompson-sampling bandit acts as the strategist, selecting command tuples from an estimate of the opponent's strategy built from in-game observations rather than opponent identity. In a controlled comparison with a flat PPO baseline trained with the same architecture, budget, curriculum and self-play league, the strategist-executor system wins significantly more often against three of the four strongest opponents on a training map, including the two strongest held-out ones (0.55 to 0.97 and 0.01 to 0.34), with no significant difference against the others.
发表机构
- NLR Royal Netherlands Aerospace Centre(荷兰皇家航空航天中心)
- Data Science Center of Excellence(数据科学卓越中心)
- Faculty of Military Sciences(军事科学学院)
- Tilburg University(蒂尔堡大学)
机构由 AI 辅助整理,请以论文原文为准。