在梦中学习,在现实中获胜:一个十英雄MOBA的连续Dyna循环
Learning in Dreams, Winning in Reality: A Continuous Dyna Loop for a Ten-Hero MOBA
浏览论文内容
中文总结 AI 辅助
本研究提出连续异步Dyna循环,在十英雄MOBA中仅于世界模型内训练策略,通过真实游戏数据在线评估与锚定,使天辉方胜率从0%提升至70.2%,揭示了模型利用与保真度继承等关键发现。
中文摘要 AI 辅助
世界模型通常从内部进行评判:通过预测损失、通过策略在想象中获得的回报,或通过其生成的帧看起来有多逼真。我们则从外部评判一个世界模型。我们学习了一个完整的十英雄MOBA(206个单位,每个英雄每个tick都行动,游戏最长可达6,000个tick)的结构化多智能体世界模型,仅在其内部以1,400个tick的自由运行想象回合训练策略,并在真实游戏中与该游戏自带的对手进行策略评估。真实游戏从不提供梯度;它提供策略自身的游戏作为世界模型的训练数据,以及一个在线评估,用于选择和锚定策略。作为连续异步Dyna循环运行,该策略在从未用于任何决策的种子上,作为天辉方赢得了70.2%的真实游戏(600局中421局;95%置信区间66.4-73.7),而仅靠梦境训练为0%,循环前为33.7%。作为夜魇方则未赢一局,游戏自带的对手自我对弈时也未赢一局。四个发现解释了这一结果。模型利用在梦境内部是不可见的:每次未锚定的运行都在几次更新内崩溃,而没有任何梦境内指标跟踪到这种崩溃。一个在其训练语料上准确的世界模型,在策略自身的游戏上却严重错误,而Dyna在这一点上进行了修复,在策略配方保持不变的情况下,真实胜率提升了+9.2个百分点。最后,策略逐机制地继承了其世界模型的保真度特征:该模型能表征宏观游戏,但不能表征群体控制、施法时机或致死性,而策略通过地图范围的压制获胜,几乎没有协调战斗。我们发布了世界模型、梦境-PPO框架、世界模型调试器、评估协议以及所有策略和日志。
英文摘要
World models are usually judged from the inside: by prediction loss, by the return a policy earns in imagination, or by how convincing their frames look. We judge one from the outside. We learn a structured, multi-agent world model of a complete ten-hero MOBA (206 units, every hero acting every tick, games of up to 6,000 ticks), train a policy only inside it with 1,400-tick free-running imagined episodes, and measure that policy in the real game against the opponent the game ships with. The real game never provides a gradient; it provides the policy's own games as training data for the world model, and an online evaluation that selects and anchors the policy. Run as a continuous asynchronous Dyna loop, the policy wins 70.2% of real games as radiant (421 of 600; 95% CI 66.4-73.7) on seeds never used for any decision, up from 0% for dream training alone and 33.7% before the loop. It wins none as dire, and neither does the shipped opponent when it plays itself. Four findings explain the result. Model exploitation is invisible from inside the dream: every unanchored run collapsed within a few updates while no in-dream metric tracked the collapse. A world model that is accurate on its training corpus is badly wrong on the policy's own games, and Dyna repairs it there, which is worth +9.2 points of real win rate with the policy recipe held fixed. Finally, the policy inherits its world model's fidelity profile mechanic by mechanic: the model represents the macro game but not crowd control, cast timing or lethality, and the policy wins by map-wide pressure with almost no coordinated fighting. We release the world model, the dream-PPO harness, a world-model debugger, the evaluation protocol, and every policy and log.