通过程序化策略搜索合成连续游戏的反应式角色行为
Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search
浏览论文内容
中文总结 AI 辅助
提出程序化策略搜索方法,通过领域特定语言和代理式草图绘制合成连续游戏的可读行为程序,在14个游戏基准上优于纯枚举和编码代理。
中文摘要 AI 辅助
我们提出了一种方法,用于将连续游戏中的反应式角色行为合成为紧凑、人类可读的程序。游戏AI实践仍然严重依赖手动编写的行为树、状态机和脚本,而学术界的强化学习通常产生不透明的神经控制器,这些控制器训练成本高且难以编辑。我们的方法通过直接在面向连续空间游戏策略的领域特定语言中进行搜索来弥合这一差距。该语言围绕反应式几何决策设计,并包含高阶构造,如方向最大化。这些构造有助于将连续行为空间离散化为可枚举的程序结构。为了使程序搜索实用化,我们引入了一大组合成反模式,这些反模式在保留行为覆盖范围的同时移除了冗余的程序形式。我们进一步将自底向上的符号枚举与来自编码代理的自顶向下指导相结合。我们提出的方法,即代理式草图绘制,让代理提出高层策略结构,并调用枚举器来完成局部程序槽。我们在一个包含14个连续游戏的基准上评估了该方法,这些游戏涵盖从经典控制任务到多智能体足球。我们发现,纯枚举通常比单独使用编码代理更高效,而组合方法则显著优于两者。我们的结果表明,程序化策略搜索可以成为游戏AI的实用创作工具:设计者指定奖励函数,系统发现可编辑的行为,这些行为有效、可移植且往往出人意料。
英文摘要
We present a method for synthesizing reactive character behaviors for continuous games as compact, human-readable programs. Game AI practice still relies heavily on manually authored behavior trees, state machines, and scripts, while academic reinforcement learning typically produces opaque neural controllers that are expensive to train and difficult to edit. Our approach bridges this gap by searching directly over a domain-specific language for continuous-space game policies. The language is designed around reactive geometric decisions and includes higher-order constructs such as direction maximization. These constructs help discretize a continuous behavior space into enumerable program structures. To make program search practical, we introduce a large set of synthesis antipatterns that remove redundant program forms while preserving behavioral coverage. We further combine bottom-up symbolic enumeration with top-down guidance from a coding agent. Our resulting method, agentic sketching, has the agent propose high-level policy structure and call an enumerator to complete local program slots. We evaluate the method on a benchmark of 14 continuous games, ranging from classic control tasks to multi-agent football. We find that pure enumeration is often more efficient than using a coding agent alone, while the combined method substantially outperforms both. Our results suggest that programmatic policy search can be a practical authoring tool for game AI: designers specify reward functions, and the system discovers editable behaviors that are effective, portable, and often surprising.
发表机构
- Brown University(布朗大学)
- Roblox(Roblox公司)
- Clemson University(克莱姆森大学)
机构由 AI 辅助整理,请以论文原文为准。