发表机构
University of Science and Technology of China; Nanjing University; Wuhan University; Meituan(中国科学技术大学; 南京大学; 武汉大学; 美团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何训练大语言模型在长期游戏中决策,提出CAST方法,利用游戏求解器状态值变化转化为求解器优势并注入RLVR,在多个游戏中优于基线,实现高平均零样本性能。
AI 中文摘要
训练大语言模型在长期游戏中行动是迈向通用决策的有前途的一步,然而基于可验证奖励的强化学习(RLVR)依赖稀疏的最终奖励,几乎无法揭示哪些决策决定成功。更密集的过程信号可以提供缺失的回合级信用,但现有来源难以同时保持低成本和准确性。我们观察到游戏求解器状态值的变化揭示了一个动作是否使状态朝着成功前进。基于此见解,我们提出了CAST(来自求解器教师的信用分配),它将这些值变化转换为求解器优势并作为回合级信号注入RLVR。我们进一步表明,在软最优求解器假设下,最大化求解器优势等同于从求解器进行策略蒸馏,只需要标量值而不是教师对数。在推箱子、扫雷和 Rush Hour 游戏中,CAST 在领域内和未见难度评估下在每场游戏中都优于所有训练的基线,并在 ALFWorld 和 WebShop 上实现了最高的平均零样本性能。我们的代码可在这个 https URL 上获取。
英文摘要
Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser process signals could supply this missing turn-level credit, but existing sources are hard to keep both cheap and accurate. We observe that changes in a game solver's state value reveal whether an action advances the state toward success. Building on this insight, we propose CAST (Credit Assignment from Solver Teachers), which converts these value changes into solver advantages and injects them into RLVR as turn-level signals. We further show that, under a soft-optimal solver assumption, maximizing the solver advantage is equivalent to on-policy distillation from the solver, requiring only scalar values rather than teacher logits. Across Sokoban, Minesweeper, and Rush Hour, CAST outperforms all trained baselines on every game under both in-domain and unseen-difficulty evaluation and achieves the highest average zero-shot performance on ALFWorld and WebShop. Our code is available at https://github.com/Wloner0809/CAST.