arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34214cs.AI

GlyphBench:语言模型强化学习的游乐场

GlyphBench: A Playground for Language-Model Reinforcement Learning

Roger Creus Castanyer, Marc-Alexandre Côté, Matthew James Sargent, Augustine N. Mavor-Parker, Glen Berseth, Pablo Samuel Castro

首次发表
浏览论文内容

中文总结 AI 辅助

GlyphBench是一个包含360多个任务的RL环境套件,通过Unicode网格观察提升智能体性能,实验表明游戏推理训练比数学或代码训练带来更强的迁移效果。

中文摘要 AI 辅助

我们推出了GlyphBench,一个用于语言模型智能体强化学习(RL)后训练的环境套件,包含超过360个涵盖多样游戏的任务。GlyphBench将空间观察渲染为二维Unicode网格,并通过一个旨在支持高效且可复现研究的统一接口连接训练、评估和轨迹回放。我们利用GlyphBench研究观察接口、推理努力和智能体框架如何影响性能,以及RL配置如何塑造学习动态。我们的结果表明,在Craftax实验中,字形观察优于原生文本和像素,并在几个BALROG环境中取得进一步提升。在100个GlyphBench任务上进行RL训练,使Qwen3.5-4B在保留的Reasoning Gym问题上得到改进,达到63.48%的准确率,并优于基础模型、数学训练基线和代码训练基线。这些实验提供了经验证据,表明来自游戏玩法的推理收益可以产生比数学或代码更强的迁移效果。总之,这些结果突显了GlyphBench作为系统研究语言模型智能体如何学习、交互和泛化的测试平台的价值。

英文摘要

We introduce GlyphBench, an environment suite for reinforcement learning (RL) post-training of language-model agents, with over 360 tasks spanning diverse games. GlyphBench renders spatial observations as two-dimensional Unicode grids and connects training, evaluation, and trajectory replay through a unified interface designed to support efficient and reproducible research. We use GlyphBench to study how observation interfaces, reasoning effort, and agent harnesses affect performance, and how RL configurations shape learning dynamics. Our results show that glyph observations outperform native text and pixels in our Craftax experiments, with further gains on several BALROG environments. RL on 100 GlyphBench tasks improves Qwen3.5-4B on held-out Reasoning Gym problems, reaching 63.48% accuracy and outperforming the base model, a math-trained baseline, and a code-trained baseline. These experiments provide empirical evidence that reasoning gains from gameplay can yield stronger transfer than math or code. Together, these results highlight GlyphBench's value as a testbed for systematic research on how language-model agents learn, interact, and generalize.

发表机构

  • Mila - Quebec Artificial Intelligence Institute(米拉-魁北克人工智能研究所)
  • Université de Montréal(蒙特利尔大学)
  • Vmax

机构由 AI 辅助整理,请以论文原文为准。

↑