发表机构
Tsinghua University; College of AI, Tsinghua University(清华大学; 清华大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究在长期游戏智能体竞赛中评估启发式学习,提出对抗性启发式学习范式,构建AAArena基准,发现Opus5.5搭配Claude Code获6金,HL在对抗性游戏具潜力但存多项挑战。
AI 中文摘要
对抗性游戏推动了从启发式搜索到强化学习的进步,但从有限样本中学习和调整策略仍然具有挑战性。AI智能体提供了一种替代方案,可将游戏经验转化为可执行策略的修订。基于启发式学习(HL),我们将对抗性启发式学习(AHL)形式化为一种范式,该范式使用AI智能体作为学习引擎来优化游戏策略和支持软件,同时保持模型权重固定。我们引入了AAArena,这是一个包含12款真实对抗性游戏和1920个存档人类程序的基准,其评估协议以现实世界游戏竞赛为模型。智能体需解读规则、选择对手、分析回放并修改游戏智能体,以在固定的比赛和评估预算内实现最高排名。我们评估了\backslash val{completedmodels}模型和工具配置:Opus5.5搭配Claude Code获得6枚金牌,而没有任何评估配置在其余6个人类排行榜中登顶。在规则规范更复杂的游戏中,性能通常较弱。进一步实验表明,对手选择和密集反馈有助于策略改进,且智能体可从自身比赛的同策略回放以及其他玩家比赛的异策略回放中学习。这些结果凸显了HL在对抗性游戏中的潜力,并指出了游戏理解、策略实施和长视域策略开发方面的持续挑战。
英文摘要
Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging. AI agents offer an alternative by turning game experience into revisions of executable policies. Building on heuristic learning (HL), we formalize Adversarial Heuristic Learning (AHL), a paradigm that uses AI agents as learning engines to refine game policies and supporting software while keeping model weights fixed. We introduce AAArena, a benchmark comprising 12 authentic adversarial games and 1,920 archived human programs, with an evaluation protocol modeled on real-world game competitions. Agents interpret rules, choose opponents, analyze replays, and revise game agents to achieve their highest ranking within fixed match and evaluation budgets. We evaluate \val{completedmodels} model and harness configurations: Opus5.5 with Claude Code earns 6 gold medals, while no evaluated configuration tops the remaining 6 human ladders. Performance is generally weaker in games with more complex rule specifications. Further experiments show that opponent selection and dense feedback support policy improvement, and that agents learn from both on-policy replays of their own matches and off-policy replays of other players' matches. These results highlight HL's potential in adversarial games and identify persistent challenges in game understanding, strategy implementation, and long-horizon policy development.