基于大型语言模型的通用游戏算法模块化发现
Modular Discovery of General Game-Playing Algorithms with Large Language Models
AI总结:
本文提出多智能体LLM元学习系统,在C++中模块化进化通用游戏搜索算法,经400+环境与15个MCTS基线对比,取得顶级评级并泛化至新游戏。
AI中文摘要:
仅依据规则在任意游戏中进行通用游戏对弈仍然具有挑战性,这源于不同游戏类别对算法要求的差异以及严格的决策时间限制。与其为特定领域手工设计搜索启发式方法,我们能否利用大型语言模型(LLMs)来发现通用游戏算法?由于语言模型能够提出并重构结构化代码,它们为探索算法设计空间提供了一个富有表现力的提案引擎。我们引入了一个多智能体LLM元学习系统,在C++中共同进化与游戏无关的程序化搜索机制,同时直接从游戏规则中合成领域启发式方法。在控制计算预算的情况下,我们在超过400个多样化环境中对发现的机制进行了基准测试,包括OpenSpiel训练和保留游戏、程序化模拟引擎,以及通过PPO训练的具有深度神经策略-价值表示的游戏。通过AlphaRank平稳分布和软孔多塞优化(SCO)与15个已建立的MCTS基线进行评估,发现的搜索机制在独立的进化运行中始终获得顶级评级,并在大多数基线上取得成对投票多数,泛化到未见的人类设计和程序化合成的游戏,并在冻结的神经网络表示上与基线保持竞争力。
英文摘要:
General Game Playing across arbitrary games from rules alone remains challenging due to differing algorithmic requirements across game classes and strict decision-time constraints. Rather than hand-designing search heuristics for specific domains, can we leverage Large Language Models (LLMs) to discover general game-playing algorithms? Because language models can propose and refactor structured code, they provide an expressive proposal engine for exploring the space of algorithmic designs. We introduce a multi-agent LLM meta-learning system to co-evolve game-agnostic procedural search mechanisms in C++ alongside domain heuristics synthesized directly from game rules. Controlling the compute budget, we benchmark the discovered mechanisms across more than 400 diverse environments, including OpenSpiel training and held-out games, procedural simulation engines, and games with deep neural policy-value representations trained via PPO. Evaluated via AlphaRank stationary distributions and Soft Condorcet Optimization (SCO) against 15 established MCTS baselines, the discovered search mechanisms consistently achieve top-tier ratings and pairwise ballot majorities over most baselines across independent evolutionary runs, generalizing to unseen human-designed and procedurally synthesized games and remaining competitive with baselines on frozen neural network representations.