面向多智能体格斗游戏中自适应NPC行为的大语言模型引导强化学习
LLM-Guided Reinforcement Learning for Adaptive NPC Behavior in Multi-Agent Combat Games
浏览论文内容
中文总结 AI 辅助
该研究提出LLM引导的RL策略选择框架,在格斗游戏中提升了NPC对平衡型对手的胜率,但对激进型对手效果不佳,同时揭示了模型策略区分能力有限的问题。
中文摘要 AI 辅助
格斗电子游戏中,基于脚本和规则的非玩家角色(NPC)通常表现出可预测的行为,经验丰富的玩家可利用这一特性;而强化学习(RL)智能体在训练后会保留固定策略,无法轻易调整策略以适应不同对手。本研究探索一种运行时策略选择框架,其中大语言模型(LLM)引导已训练的RL策略,且不修改其底层行为。为此,我们在Unity中训练5个共享PPO策略的NPC智能体,将策略独立行动的基线配置与增强配置对比:增强配置通过Ollama访问本地部署的Mistral 7B模型,每5秒读取一次实时游戏状态并分配4种战术标签之一。我们在600局游戏中针对3种脚本化对手类型评估两种配置,并采用Mann-Whitney U检验分析结果。针对一局游戏中会改变战术的平衡型对手,LLM增强型智能体的胜率从11%提升至24%,实现翻倍以上,且游戏时长显著更长;针对规避型对手,增强型智能体胜率更高且击杀速度更快,尽管其较短的游戏时长不符合严格假设定义;针对激进型对手,LLM近乎持续偏好包围的策略适得其反。对2430次策略选择的分析显示,无论对手类型如何,包围策略的选择占比达83.8%,表明该模型规模下的零样本策略区分能力有限。这些结果证明了LLM引导的运行时策略选择在自适应多智能体游戏AI中的潜力与局限性。
英文摘要
Scripted and rule-based non-player characters (NPCs) in combat video games often exhibit predictable behaviors that experienced players can exploit, while reinforcement learning (RL) agents typically retain a fixed policy after training and cannot readily adapt their strategy to different opponents. We investigate a runtime strategy-selection framework in which a large language model (LLM) guides a trained RL policy without modifying its underlying behavior. To demonstrate this, we train five NPC agents with a shared PPO policy in Unity and compare a baseline configuration, in which the policy acts independently, with an LLM-augmented configuration in which a locally hosted Mistral 7B model, accessed through Ollama, reads the live game state every five seconds and assigns one of four tactical tags. We evaluate both configurations against three scripted opponent types across 600 episodes and analyze outcomes using the Mann-Whitney U test. Against a Balanced opponent that changes tactics during an episode, the LLM-augmented agents more than doubled their win rate from 11% to 24% and produced significantly longer episodes. Against an Evasive opponent, the augmented agents achieved a higher win rate and faster kills, although their shorter episode duration did not satisfy the strict hypothesis definition. Against an Aggressive opponent, the LLM's near-constant preference for encirclement was counterproductive. Analysis of 2,430 strategy selections showed that Surround was selected in 83.8% of cases regardless of opponent type, indicating limited zero-shot strategic differentiation at this model scale. These results demonstrate both the potential and limitations of LLM-guided runtime strategy selection for adaptive multi-agent game AI.
发表机构
- Heriot-Watt University(赫瑞瓦特大学)
机构由 AI 辅助整理,请以论文原文为准。