发表机构
University of California, Berkeley; Johns Hopkins University(加州大学伯克利分校; 约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究部分可观测动态博弈中未知对手的K级策略编排问题,提出基于强化学习的编排器RLBO,实验表明其优于分类式编排器,且训练效率高,支持将策略层级作为编排资源。
AI 中文摘要
K级推理生成一个针对不同推理级别对手的专业化策略层级。当对手的级别未知时,常见的部署规则会估计该级别并选择相应的响应。在动态、部分可观测的博弈中,这种选择会重复进行,每次选择都会影响后续的状态和观测。与最可能的对手级别相关联的响应不一定能从当前历史中最大化期望回报。我们将此部署问题形式化为在部分可观测马尔可夫博弈中对固定、预训练的策略库进行动态编排。我们比较了使用离线数据或在线策略数据聚合训练的分类式编排器(CBOs)与旨在最大化期望折扣回报的基于强化学习的编排器(RLBO)。在追逃实验中,在线策略训练改善了分类和回报,但RLBO获得的回报高于在线策略和离线CBOs。给定预训练库,RLBO还达到了与直接在追捕者动作空间上训练的策略相当的性能,且训练时间步更少。这些发现支持将K级策略层级视为编排的资源,而非部署的处方。
英文摘要
Level-$K$ reasoning generates a hierarchy of policies specialized to opponents with different reasoning levels. When an opponent's level is unknown, a common deployment rule estimates that level and selects the corresponding response. In a dynamic, partially-observed game, this selection is repeated, with each choice shaping subsequent states and observations. The response associated with the most likely opponent level need not maximize expected return from the current history. We formulate this deployment problem as dynamic orchestration of a fixed, pretrained policy library in a partially observable Markov game. We compare classification-based orchestrators (CBOs) trained using offline data or on-policy data aggregation with a reinforcement-learning-based orchestrator (RLBO) trained to maximize expected discounted return. In pursuit-evasion experiments, on-policy training improves classification and return, yet RLBO achieves higher return than the on-policy and offline CBOs. Given a pretrained library, RLBO also reaches performance comparable to a policy trained directly over the pursuers' action space with fewer training timesteps. These findings support treating a hierarchy of level-$K$ policies as a resource for orchestration, not a prescription for deployment.