发表机构
Frisson Labs(Frisson Labs)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Faynt提出10M和75M参数的Transformer策略,通过强化学习、蒸馏和课程训练,在《任天堂明星大乱斗DX》中实现对26个角色的统一控制,以98.4%胜率击败专精对手,并开源权重与基准。
AI 中文摘要
我们提出了Faynt,一个包含1000万和7500万参数Transformer策略的系列,用于《任天堂明星大乱斗DX》,每个策略通过单个检查点控制全部26个角色。经过强化学习(RL)后,1000万参数模型在240场同角色对局中赢得240场(98.4%),对手为十四个专精和多角色版本,在其支持的角色名单上,且对每个版本均保持胜绩。这些对手保留21帧或24帧的动作延迟;Faynt不使用额外延迟,我们尚未隔离此差异的影响。在另一项针对私下提供的零延迟Slippi-AI模型的评估中,1000万参数模型在两个条件设置下赢得全部68场对局。我们研究架构、优化、扩展和超参数迁移,以指导在约84万个人类回放上的预训练。训练后结合基于排名和基于结果的课程、7500万到1000万的蒸馏,以及仅限于Fox镜像对局的RL。在初始152场基准测试中,监督训练的1000万模型赢得69.7%的对局,而预训练的7500万模型为45.4%,尽管其整体保留的控制器预测损失更高。用于监督检查点选择的加权验证损失与所有四个预训练和监督策略的胜率排序一致。在监督训练后,两个模型每分钟受到的伤害更少,建立更大的早期领先优势,并在失去第一条命后更频繁获胜。在记录的游戏状态上进行优化推理,在NVIDIA T4上,1000万模型平均每次决策5.2毫秒,7500万模型为8.7毫秒,不包括模拟器执行和通信。我们开源了权重、两个基准测试套件以及一个用于自动化模型锦标赛的平台。
英文摘要
We introduce Faynt, a family of 10M- and 75M-parameter Transformer policies for Super Smash Bros. Melee, each controlling all 26 characters with a single checkpoint. After reinforcement learning (RL), the 10M wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character releases on their supported rosters, with a winning record against every release. These opponents retain 21- or 24-frame action delays; Faynt uses no added delay, and we have not isolated the effect of this difference. In a separate evaluation against a privately supplied zero-delay Slippi-AI model, the 10M wins all 68 games across two conditioning settings. We study architecture, optimization, scaling, and hyperparameter transfer to guide pretraining on approximately 840,000 human replays. Post-training combines rank- and outcome-based curricula, 75M-to-10M distillation, and RL restricted to Fox mirror matches. On the initial 152-game benchmark, the supervised 10M wins 69.7% of games, compared with 45.4% for the pretrained 75M, despite higher overall held-out controller-prediction loss. The weighted validation loss used for supervised checkpoint selection agrees with the win-rate ordering of all four pretrained and supervised policies. After supervised post-training, both models take less damage per minute, build larger early leads, and win more often after losing the first life. Optimized inference on recorded game states averages 5.2 ms per decision for the 10M and 8.7 ms for the 75M on an NVIDIA T4, excluding emulator execution and communication. We open-source the weights, both benchmark suites, and a platform for automated model tournaments.
Comments54 pages. Preprint, in review