发表机构
New York University; University of Washington; University of Southern California; Columbia University; Dimension Gate(纽约大学; 华盛顿大学; 南加州大学; 哥伦比亚大学; 维度之门)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究多工具RL在编码智能体中的信用分配与可迁移性,发现评估工具集对性能影响远大于训练分组规则,跨工具集信用未提升能力可迁移性,建议多工具RL报告需说明分组边界并在未见过的工具集测试。
AI 中文摘要
智能体强化学习(RL)越来越多地通过完整执行工具集(harness)运行,而多工具方法包含两种选择:让策略暴露于多个工具集,以及在一个相对优势组内比较它们的奖励。我们在仓库级代码任务中单独研究第二种选择。从Qwen3-8B监督学习的预训练起点出发,我们重复使用来自Aider、OpenHands、Qwen Code和SWE-agent的相同冻结任务-工具集记录,在相同的更新次数下,采用两种组相对策略优化(GRPO)规则:Within(每个任务-工具集对为一个组)和Cross(同一任务内的工具集合并),并使用密封的SWE-bench Verified评估器在四个源工具集和一个训练时未见过的最小工具集上对每个检查点进行评分。评估工具集是主要变量:在24000次密封评估中,它将平均解决率从2.14%提升至9.27%,提升了4.3倍,而训练方法仅将其提升了1.16倍。分组规则则不是主要因素:在未见过的工具集上,Cross规则减去Within规则的差值为+0.25个百分点,95%置信区间为[-0.48, +1.02],在每个任务尝试8次时;在三个训练种子的合并结果中,差值为+0.16个百分点,95%置信区间为[-0.41, +0.72],而单个种子的估计值甚至出现符号反转。每个规则自身的种子范围(0.42至0.45个百分点)超过了两者之间的差值。两种规则的最大增益都出现在相同的源工具集上。合并的优势能够反映工具集:折衷分类器从Cross规则的优势中以高于打乱标签基线4.48个百分点的幅度恢复了生成工具集,而从Within规则中则无法恢复,并且两种规则在每个工具集内仍达到相同的未见过的得分和动作分布。重新收集一半的在线策略训练数据不会改变这一结果。跨工具集信用分配会产生配置适应性,但不会比工具集内的信用分配带来更可迁移的能力。多工具RL的报告应说明分组边界并在未见过的工具集下进行测试。
英文摘要
Agent reinforcement learning (RL) increasingly runs through full execution harnesses, and a multi-harness recipe mixes two choices: exposing the policy to several harnesses, and comparing their rewards inside one relative-advantage group. We isolate the second choice in repository-level coding. From one Qwen3-8B supervised warm start we replay the same frozen task-harness records from Aider, OpenHands, Qwen Code, and SWE-agent, with the same number of updates, under two rules for group-relative policy optimization (GRPO), Within (one group per task-harness pair) and Cross (harnesses pooled within a task), and score every checkpoint with a sealed SWE-bench Verified oracle on four source harnesses and a minimal harness held out of training. The evaluation harness is the dominant variable: across 24,000 sealed evaluations it moves the mean solve rate from 2.14\% to 9.27\%, a factor of 4.3, where the training recipe moves it by 1.16. The grouping rule is not. On the held-out harness, Cross minus Within is +0.25 pp, 95\% confidence interval [-0.48, +1.02], at eight attempts per task, and +0.16 [-0.41, +0.72] pooled over three training seeds whose individual estimates change sign. Each rule's own seed range, 0.42 to 0.45 pp, exceeds the difference between them. Both rules place their largest gains on the same source harness. The pooled advantage carries the harness: an out-of-fold classifier recovers the generating harness from Cross's advantage +4.48 pp above the shuffled-label baseline and from Within's not at all, and the two rules still reach the same held-out score and action distribution inside each harness. Re-collecting half the training data on-policy does not change this. Cross-harness credit yields configuration adaptation and no more portable capability than within-harness credit. Multi-harness RL reports should state the grouping boundary and test under an unseen harness.
Comments21 pages, 1 figure, 10 tables. Preprint