可指导的交互式游戏智能体
Coachable agents for interactive gameplay
浏览论文内容
中文总结 AI 辅助
提出结合通用价值函数近似器与精心选择的训练场景、学习算法和数据增强的框架,使智能体在复杂领域(如《地平线:西之绝境》、《GT赛车》和类人机器人)中实时展现不同风格,同时保持主任务性能。
中文摘要 AI 辅助
强化学习已被证明是创建先进AI和机器人系统的宝贵工具,贡献从游戏到机器人再到基础模型的各个方面。通过试错,这些AI系统通常学习一种接近最优的行为来解决任务。然而,在许多用例中,人们希望在一定程度上实时控制任务解决的方式。我们将这些对核心任务的修改称为风格。我们结合通用价值函数近似器(UVFA)与精心选择的训练场景、学习算法和数据增强,创建了一个框架,用于指导智能体在复杂领域中展现风格。我们展示了该框架在AAA视频游戏《地平线:西之绝境》和《GT赛车》以及开源类人机器人测试领域的应用。尽管领域性质不同——赛车、风格化游戏战斗和类人机器人行走——每个智能体都表现出对风格请求的强一致性,同时仍满足其领域的主要任务。重要的是,本文概述的技术允许最终用户在运行时选择最终行为,从而灵活控制最终执行的性能。
英文摘要
Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from game playing to robotics to foundation models. Through trial-and-error, these AI systems typically learn one, near-optimal behavior to solve their tasks. However, there are many use cases in which one would like to assert some level of control, preferably in real time, over how the task is solved. We refer to these modifications of a core task as styles. We combine universal value function approximators (UVFAs) with carefully selected training scenarios, learning algorithms, and data augmentation to create a framework for coaching agents that exhibit styles in complex domains. We demonstrate the framework's application in the AAA video games Horizon Forbidden West and Gran Turismo, and in an open-source humanoid test domain. Despite the different nature of the domains -- car racing, stylized game combat, and humanoid walking -- each agent shows strong coherence to the style requests while still satisfying the main task in its domain. Importantly, the techniques outlined in this paper allow an end user to choose the final behavior at run time, giving them flexible control over the final executed performance.
发表机构
- Sony AI, Zurich, Switzerland(索尼AI,苏黎世,瑞士)
- Sony AI, North America (various locations)(索尼AI,北美(多地))
- Sony AI, Tokyo, Japan(索尼AI,东京,日本)
机构由 AI 辅助整理,请以论文原文为准。