发表机构
Surge AI(Surge AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对长视野智能体,通过两阶段SFT-再-RL流程后训练Qwen3.5-122B-A10B模型,在5项外部基准实现性能提升,证明长视野多工具后训练的工作方式可跨领域迁移。
AI 中文摘要
对于自包含环境中的强化学习(RL),策略可通过利用环境特定规律(工具模式、 grader解析、任务模板)获取奖励,而非获得可迁移技能,分布内保留集也会共享这些规律。我们认为关键问题是行为层面,即训练后的智能体如何行动,而跨基准迁移是研究该问题的合适方向。我们在363个跨27类的长视野模型上下文协议(MCP)任务上,对开放权重混合专家模型(Qwen3.5-122B-A10B)进行后训练,采用两阶段SFT-再-RL流程。Toolathlon性能为初始基础家族和SFT教师选择提供了依据,但训练中未使用任何外部基准任务或grader,也未将外部分数用于奖励、训练超参数、训练检查点选择或停止条件。在贪婪pass@1指标下,该训练模型在5项外部评估中均优于基础模型:Toolathlon(+9.6个百分点)、τ²-Bench(+5.3个百分点)、BFCL-V4(+3.5个百分点)、SWE-Bench Pro(+5.8个百分点)和Terminal-Bench 2(+2.8个百分点)。尽管训练集不含软件工程任务,但两项软件工程基准均实现了性能提升。探索性成对轨迹分析发现了四种反复出现的行为差异:更谨慎的局部目标形成、构建与目标相关的工作状态、在局部修复中保持父目标稳定,以及验证完成情况,这些差异在办公工作流和代码中以类似形式出现。这些结果提供了描述性证据,表明长视野多工具后训练可改变工作方式,且该工作方式可迁移至训练领域之外。
英文摘要
For reinforcement learning (RL) in self-contained environments, a policy can get rewards by exploiting environment-specific regularities (tool schemas, grader parsing, task templates) rather than by acquiring transferable skill, and an in-distribution holdout shares those regularities. We argue that the discriminating question is behavioral, namely how a trained agent acts, and that cross-benchmark transfer is the right place to look for it. We post-train an open-weight mixture-of-experts model (Qwen3.5-122B-A10B) on 363 long-horizon Model Context Protocol (MCP) tasks across 27 categories, using a two-stage SFT-then-RL pipeline. Toolathlon performance informed the initial base-family and SFT-teacher choices, but no external-benchmark task or grader entered training and no external score informed the reward, training hyperparameters, trained-checkpoint selection, or stopping. At greedy pass@1, the trained model improves over the base on five reported external evaluations: Toolathlon (+9.6 pp), $τ^2$-Bench (+5.3 pp), BFCL-V4 (+3.5 pp), SWE-Bench Pro (+5.8 pp), and Terminal-Bench 2 (+2.8 pp). Both software-engineering benchmarks improve despite the training collection containing no software-engineering tasks. An exploratory paired-trajectory analysis identifies four recurring behavioral differences (more careful local-goal formation, building goal-relevant working state, keeping parent goals stable through local repairs, and verifying completion) that appear in analogous forms across office workflows and code. These results provide descriptive evidence that long-horizon multi-tool post-training can change ways of working that transfer beyond its training domain.
CommentsAccepted at the COLM 2026 Workshop on Agent Behavior. 11 pages, 4 tables