arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38881cs.AI

STRATA:面向实时策略游戏的角色对齐分层智能体的自学习

STRATA: Self-Learning Through Role-Aligned Tiered Agents for Real-Time Strategy Games

Xinhe Tian, Xiaoyue Zhang, Ziyou Zhang, Jiacheng Li, Xiaoqiang Jin, Qianchuan Zhao, Gaochen Cui

首次发表
浏览论文内容

中文总结 AI 辅助

STRATA提出角色对齐的分层智能体系统,通过跨游戏自学习生成经验卡片,将《红色警戒》中的胜率从30%提升至100%,并适应不同对手风格。

中文摘要 AI 辅助

实时策略(RTS)游戏要求智能体在长时间对局中协调经济发展、生产与建设、基地防御、单位编组以及进攻时机。已有研究将大语言模型应用于RTS游戏的指挥决策,使智能体能够读取文本化游戏状态并生成高层计划。然而,长推理延迟可能导致它们错过关键战术事件。完整RTS对局的复杂性和战术多样性也使现有系统严重依赖人工编写的基于经验的提示,从过往游戏中持续学习的能力有限。我们提出STRATA,一个用于《红色警戒》的角色对齐分层系统,具备跨游戏自学习能力。STRATA分别将游戏中的战略、后勤和战术决策分配给战略智能体(SA)、后勤智能体(LA)和战术智能体(TA)。SA基于全局游戏状态和相关经验卡片生成高层指令,而LA和TA负责后勤和战术执行。每场对局后,评审智能体(RA)从游戏轨迹中提取候选经验,利用后续对局的证据进行验证和修订,并将多场游戏支持的战略经验压缩为简洁的经验卡片供SA检索。我们通过经验卡片的形成、学习前后的完整对局比较以及针对不同风格AI对手的经验学习来评估STRATA。在固定场景下,使用学习到的经验卡片将观察到的胜率从30%提升至100%。针对不同风格AI对手的序列学习也产生了不同的长期战略经验。

英文摘要

Real-time strategy (RTS) games require agents to coordinate economic development, production and construction, base defense, unit organization, and attack timing over long matches. Existing studies have applied large language models to command decision-making in RTS games, enabling agents to read textual game states and generate high-level plans. However, long inference latency can cause them to miss critical tactical events. The complexity and tactical diversity of full RTS matches also leave existing systems heavily dependent on manually written experience-based prompts, with limited ability to learn continuously from past games. We present STRATA, a role-aligned hierarchical system with cross-game self-learning for Red Alert. STRATA assigns in-game strategic, logistical, and tactical decisions to a Strategic Agent (SA), Logistics Agent (LA), and Tactical Agent (TA), respectively. The SA generates high-level directives based on the global game state and relevant experience cards, while the LA and TA handle logistics and tactical execution. After each match, a Review Agent (RA) derives candidate experience from game traces, validates and revises it using evidence from subsequent matches, and compresses strategic experience supported across multiple games into concise experience cards for SA retrieval. We evaluate STRATA through the formation of experience cards, full-match comparisons before and after learning, and experience learning against AI opponents with different play styles. Under a fixed scenario, using the learned experience cards increases the observed win rate from 30% to 100%. Sequential learning against AI opponents with different play styles also produces distinct long-term strategic experience.

发表机构

  • Beijing University of Aeronautics and Astronautics(北京航空航天大学)
  • Tsinghua University(清华大学)
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑