arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PlayTrain:一种用于LLM生成的可适配JavaScript游戏的高效强化学习框架

PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games

Ryan Truong, Lance Ying, Samuel J. Gershman, Kazuki Irie

arXiv 2609.09059首次发表:更新:

发表机构

Harvard University; MIT; Kempner Institute for the Study of Natural and Artificial Intelligence; Yale University(哈佛大学; 麻省理工学院; 肯普纳自然与人工智能研究所; 耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PlayTrain利用大语言模型生成JavaScript游戏,并在标准gym环境中高效训练强化学习智能体,支持克隆和修改Atari及ProcGen游戏,实现每秒超百万次决策,简化VGE开发流程。

AI 中文摘要

虽然许多视频游戏环境(VGEs)在推进强化学习(RL)方面发挥了关键作用,但开发新的VGE或修改现有VGE以支持新功能,一直是一个需要大量手工编码的费力过程。在此,我们提出PlayTrain,一个强化学习框架,它结合了大型语言模型(LLMs)从最小的人类提示中稳健地生成JavaScript(JS)游戏的能力,以及一个能在标准'gym'环境中运行任何JS游戏的高效流水线。不仅最近的LLM特别擅长编写JS代码,而且JS格式还允许用户轻松游玩生成的VGE,而PlayTrain使我们能够在完全相同的游戏上训练RL智能体。我们展示了PlayTrain的多种用例,包括用简单JS克隆知名的Atari和ProcGen游戏,其中PlayTrain在单个GPU节点上以每秒超过100万次智能体决策的速度端到端训练基于像素的智能体;以及创建其修改版本(例如,支持新颖的测试集、程序化生成逻辑或游戏动态)。通过PlayTrain,我们重新构想了RL VGE的开发:我们所需要的只是一个通过LLM生成和修改的单一JS文件。我们讨论了PlayTrain解锁的有前景的未来RL研究方向。

英文摘要

While many video-game environments (VGEs) have played crucial roles in advancing reinforcement learning (RL), developing novel VGEs or modifying existing ones to support new features, has been a laborious process requiring extensive hand-coding. Here we present PlayTrain, an RL framework that combines the abilities of large language models (LLMs) to robustly generate JavaScript (JS) games from a minimal human prompt, and an efficient pipeline that can run any JS game in a standard 'gym' environment. Not only are recent LLMs particularly good at writing JS code, but the JS format also allows users to easily play generated VGEs, while PlayTrain enables us to train RL agents on the exact same games. We demonstrate multiple use cases of PlayTrain, including cloning well-known Atari and ProcGen games in simple JS, where PlayTrain trains pixel-based agents end-to-end at over 1M agent-decisions per second on a single GPU node; and creating modified versions thereof (e.g., that support novel test sets, procedural generation logics, or game dynamics). Through PlayTrain, we reimagine RL VGE development: all we need is a single JS file, generated and modified through an LLM. We discuss promising future RL research directions that PlayTrain unlocks.

Comments36 pages. v2: adds a Craftax-Classic figure and AI-use, ethics and reproducibility statements, removes the bigfish code figure, condenses the main text, and revises the appendix. Code and data: https://github.com/heyodog0/playtrain/releases/tag/paper-v2

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑