发表机构
University of Washington; Stanford University; Northeastern University; Carnegie Mellon University; Massachusetts Institute of Technology; National University of Singapore; Seoul National University; Stevens Institute of Technology; University of Chicago(华盛顿大学; 斯坦福大学; 东北大学; 卡内基梅隆大学; 麻省理工学院; 新加坡国立大学; 首尔大学; 史蒂文斯理工学院; 芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SPADE是一种双角色自博弈强化学习框架,让单个LLM同时担任环境设计者与推理智能体,在多类基准上显著提升了语言智能体的性能,推动了开放式自我改进。
AI 中文摘要
持续的自我改进需要不断扩展的、由自身生成的、多样化的、自适应的目标池。对于语言智能体而言,现有的训练环境池(人工策划、静态合成或冻结验证器)在学习者规模扩大时,目标分布保持固定。我们引入SPADE(Self-Play in Adaptive Synthetic Executable Environments,自适应合成可执行环境中的自博弈),这是一种自博弈强化学习框架,其中单个大语言模型(LLM)扮演两个角色:环境设计者,编写完整的、长程的训练环境作为可执行代码,具备OpenAI Gym风格的reset()/step()接口;以及推理智能体,学习在这些环境中执行动作。每个角色都是有状态的多轮环境(包含状态转换、奖励函数和验证代码),因此一个接口可覆盖推理问题和多步骤智能体工具使用。推理智能体的遗憾通过其在有特权提示和无特权提示下的奖励差距来估计;在优化该遗憾信号时,环境设计者学会将环境目标设定在智能体能力的边缘,同时保持其可行性。通过大量实验,我们发现几个对成功至关重要的组件:将环境设计者基于从大型预训练语料库中采样的文档进行锚定,并为其提供累积的环境记忆。扩展到300亿参数模型时,SPADE在8个保留的数学、科学、代码和推理基准上,比最强的固定环境基线平均提升5.3;在BFCL-v4多轮工具使用设置上提升5.7,在ACEBench-Agent上提升13.9;在游戏设置中,与最强基线的差距随模型规模增大而扩大。通过将环境设计本身变为可学习组件,SPADE为开放式自我改进迈出了具体的一步。
英文摘要
Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales. We introduce SPADE (Self-Play in Adaptive Synthetic Executable Environments), a self-play RL framework in which a single LLM plays two roles: an Environment Designer that writes complete, long-horizon training environments as executable code with an OpenAI Gym-style reset()/step() interface, and a Reasoning Agent that learns to act in them. Each is a stateful, multi-turn environment (state transitions, reward functions, and verification code), so one interface spans reasoning problems and multi-step agentic tool use. The Reasoning Agent's regret is estimated using the gap between its reward with and without privileged hints; in optimizing this regret signal the Environment Designer learns to target environments at the edge of the agent's capabilities while keeping them feasible. Through extensive experimentation, we find several components critical to success: grounding the Environment Designer on documents sampled from a large pretraining corpus, and giving it an accumulated environment memory. Scaling to 30B-parameter models, SPADE improves over the strongest fixed-environment baseline by +5.3 on average across eight held-out math, science, code, and reasoning benchmarks, and lifts the tool-use setting by +5.7 on BFCL-v4 multi-turn and +13.9 on ACEBench-Agent; on the games setting, the margin over the strongest baseline grows with model scale. By making environment design itself a learnable component, SPADE takes a concrete step toward open-ended self-improvement.
CommentsWork in progress. Project page: https://spade-rl.github.io ; Code: https://github.com/spade-rl/spade