发表机构
Surge AI(Surge AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过在办公工作流程的长程多工具智能体任务上后训练Qwen3.5-122B-A10B,发现模型在无相关训练的软件工程基准SWE-Bench Pro上性能提升,证实长程后训练可跨域强化目标导向执行行为。
AI 中文摘要
长程任务要求智能体在嵌套及分支工作中维持连贯的状态与目标,我们将该能力称为目标导向执行(GDE),即反复应用四种行为:选择目标、构建任务相关状态、保持对高层目标的忠实度,以及对照环境验证完成情况。我们假设长程后训练可在各领域强化这些行为,为此我们在源自办公工作流程的363个长程多工具智能体(LHMTA)任务上对Qwen3.5-122B-A10B进行后训练。该数据集不含软件工程任务,但模型在SWE-Bench Pro上的pass@1提升了5.8个百分点。匹配轨迹分析显示,模型在办公工作流程和软件仓库的所有四种GDE行为上均有提升;SWE-Bench Pro的整体统计数据表明,信息收集、实现及验证环节也出现了相关变化。综合来看,这些结果支持一种行为层面的解释:长程后训练改变了模型组织与应用跨任务知识的方式,其影响超出了训练领域的范围。
英文摘要
Long-horizon tasks require agents to maintain coherent state and goals across nested and branching work. We call this capability goal-directed execution (GDE): the repeated application of four behaviors, namely selecting goals, constructing task-relevant state, maintaining fidelity to higher-level objectives, and verifying completion against the environment. We hypothesize that long-horizon post-training strengthens these behaviors across domains. We test this by post-training Qwen3.5-122B-A10B on 363 Long-Horizon Multi-Tool Agent (LHMTA) tasks drawn from office workflows. The collection contained no software-engineering tasks, yet the model's pass@1 improved by 5.8 points on SWE-Bench Pro. Matched trajectory analysis shows gains in all four GDE behaviors in both office workflows and software repositories. Aggregate SWE-Bench Pro statistics showed related changes in information gathering, implementation, and verification. Together, the results support a behavioral interpretation in which long-horizon post-training changed how the model organized and applied knowledge across tasks, with effects extending beyond the training domain.
Comments20 pages, 8 figures, 5 tables