Envs-FORGE:面向智能体强化学习的前沿优化、奖励驱动的环境合成方法
Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis for Agent RL
- IDEA Research(IDEA研究院)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- DataArcTech Ltd.(DataArcTech有限公司)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
Envs-FORGE是面向智能体RL的奖励驱动环境合成方法,通过MILP选动作生成训练环境,在多模型及数据集上显著提升Pass@1等指标。
中文摘要 AI 辅助
终端智能体的强化学习(RL)需要具备可靠奖励和合适难度的可执行训练环境。Few-shot、Self-Instruct、Evol-Instruct等固定方案对每个样本(seed)采用相同的提示策略,即便当前策略需要更难、更简单或完全不同的任务也不调整。本文提出Envs-FORGE,一种将验证器奖励转换为每个样本环境合成动作的提示策略。Envs-FORGE估计样本通过率,围绕目标学习前沿对六个投影方向动作评分,并求解每个样本的混合整数线性规划(MILP)以选择条件生成的动作。所选动作驱动指令、装置、oracle解、测试及Docker环境的同步重写;仅经gold验证的环境束进入RL训练。索引式MILP形式还支持投资组合规划的可选软技能覆盖。在Qwen 3.5 35B模型上,Envs-FORGE在tb-core数据集上将Pass@1指标较基线提升9.2个百分点(从40.0%升至49.2%),在tb-2.0数据集上提升6.4个百分点(从23.0%升至29.4%),比最强固定方案基线分别超出2.4和2.1个百分点;在SWE-bench Verified上达到77.1%,而基线为73.4%;在4B至35B规模的评估模型中,tb-core的提升幅度为6.8至9.2个百分点。所有合成方法均导出100个验证环境,使用227万至288万合成token,使对比处于相同下游训练集规模和操作尺度。源代码可在指定URL获取。
英文摘要
Reinforcement learning (RL) for terminal agents needs executable training environments with reliable rewards and useful difficulty. Fixed recipes such as few-shot, Self-Instruct, and Evol-Instruct apply the same prompting policy to every seed, even when the current policy would benefit from a harder, easier, or simply different task. We present Envs-FORGE, a prompting policy that converts verifier rewards into per-seed environment-synthesis actions. Envs-FORGE estimates seed pass rates, scores six projection--direction actions around a target learning frontier, and solves a per-seed mixed-integer linear program (MILP) to choose the action that conditions generation. The selected action drives synchronized rewriting of the instruction, fixtures, oracle solution, tests, and Docker environment; only gold-verified bundles enter RL training. The indexed MILP form also supports optional soft skill coverage for portfolio planning. On Qwen 3.5 35B, Envs-FORGE improves Pass@1 over Base by 9.2 percentage points on tb-core (40.0% to 49.2%) and 6.4 points on tb-2.0 (23.0% to 29.4%), exceeding the strongest fixed-recipe baseline by 2.4 and 2.1 points. It reaches 77.1% on SWE-bench Verified versus 73.4% for Base, and improves tb-core by 6.8--9.2 points across the evaluated 4B--35B models. All synthesis methods export 100 verified environments and use 2.27M--2.88M synthesis tokens, placing the comparison at the same downstream training-set size and the same operational scale. The source code is available at https://github.com/DataArcTech/DataArc-SynData-Toolkit/.