发表机构
Czech Technical University in Prague; Carnegie Mellon University; Artificial Intelligence Center; Strategy Robot, Inc.; Strategic Machine, Inc.; Optimized Markets, Inc.(布拉格捷克理工大学; 卡内基梅隆大学; 人工智能中心; 策略机器人公司; 战略机器公司; 优化市场公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对两人零和非完美信息博弈,提出将小工具博弈扩展至强化学习场景的策略梯度算法,通过隐式表示消除子博弈规模限制,经实验验证可提升策略性能。
AI 中文摘要
测试时推理已显著提升了从游戏到语言模型等领域的性能。然而,在两人零和非完美信息博弈中,对所得策略性能具有形式保证的测试时策略调整仍是一项挑战。现有解决方案仅限于表格方法或单梯度步更新。本研究将策略梯度算法作为可扩展测试时推理的方法进行探究,把用于测试时搜索的表格技术“小工具博弈(gadget game)”扩展至强化学习场景。与先前方法不同,我们通过改进采样和神经策略隐式表示小工具博弈,而非显式构建它,从而消除了子博弈规模的限制。此外,我们形式化证明,与先前的表格算法不同,正则化策略梯度算法即使不依赖小工具博弈,也能限制测试时推理导致的可能策略性能下降。我们在小型和大型博弈中的评估证实,额外的测试时训练通常会相对于蓝图策略大幅提升性能。
英文摘要
Test-time reasoning has significantly improved performance in domains ranging from games to language models. However, test-time policy changes with formal guarantees on the performance of the resulting strategy remain a challenge in two-player zero-sum imperfect-information games. Existing solutions are limited to tabular methods or single gradient step updates. In this work, we investigate policy-gradient algorithms as a method for scalable test-time reasoning. We extend the concept of gadget game, tabular technique for test-time search, to the reinforcement learning setting. Unlike prior approaches, we represent the gadget game implicitly by modified sampling and neural policy rather then explicitly by constructing it, thereby removing constraints on subgame size. Furthermore, we formally prove that, unlike prior tabular algorithms, regularized policy-gradient algorithms limit possible strategy degradation caused by test-time reasoning, even without the gadget games. Our evaluation across small- and large-scale games confirms that additional test-time training often substantially improves performance relative to the blueprint strategy.