WEFT:面向通用智能体的工具使用后训练规模化
WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents
浏览论文内容
中文总结 AI 辅助
提出WEFT,通过全系统演化构建、执行驱动自我演化与稳定后训练,扩展工具使用后训练,显著提升通用智能体在多个基准上的性能。
中文摘要 AI 辅助
近期扩展工具使用后训练的工作主要集中于合成可执行环境,而环境仅构成更广泛的智能体交互系统(包括环境、任务、智能体框架和评估器)的一个组成部分。然而,孤立地扩展环境并不能保证模型性能获得相应的提升,因为可靠的学习信号依赖于智能体交互系统中所有组件之间的连贯交互。为解决这一问题,我们提出了WEFT(Whole-system Evolution For Tool-use Post-training,工具使用后训练的全系统演化),它将可扩展的智能体交互系统构建、执行驱动的自我演化与稳定的后训练相结合。WEFT在环境广度、任务复杂度和交互多样性三个维度上扩展智能体交互系统的构建。执行驱动的自我演化迭代地利用执行轨迹和状态证据来归因失败并修正相关组件,通过新的回滚评估这些更改,并为后续演化轮次提供证据。为了实现大规模稳定后训练,WEFT同时解决了优化可靠性和执行可靠性两个问题:前缀保留采样保留了已验证的进展,原子回合信用分配将学习信号局部化;而MegaMCP在共享工具服务上的并发回滚中维持隔离且可恢复的状态。跨多种模型和基准的大量实验证明了WEFT在工具使用后训练中的有效性。WEFT-8B和WEFT-14B在BFCL V4、τ²-Bench和Claw-Eval上均优于所有评估过的同规模环境扩展基线。特别地,WEFT-14B相比Agent-World-14B分别提升了6.41、2.23和12.27个百分点。WEFT-35B-A3B进一步将这些增益扩展到更具挑战性的长时程工作流基准,包括Toolathlon-Verified和AutomationBench。
英文摘要
Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee commensurate gains in model performance, because reliable learning signals depend on coherent interactions among all components of the agentic interaction system. To address this problem, we introduce WEFT (Whole-system Evolution For Tool-use Post-training), which couples scalable agentic interaction system construction, execution-driven self-evolution, and stable post-training. WEFT scales agentic interaction system construction across environment breadth, task complexity, and interaction diversity. Execution-driven self-evolution iteratively uses execution traces and state evidence to attribute failures and revise the responsible components, with fresh rollouts evaluating the changes and providing evidence for subsequent evolution rounds. For stable post-training at scale, WEFT addresses both optimization and execution reliability: prefix-preserving sampling retains verified progress and atomic-turn credit assignment localizes learning signals, while MegaMCP maintains isolated, recoverable state across concurrent rollouts over shared tool services. Extensive experiments across various models and benchmarks demonstrate the effectiveness of WEFT for tool-use post-training. WEFT-8B and WEFT-14B outperform all evaluated matched-size environment-scaling baselines on BFCL V4, $τ^2$-Bench, and Claw-Eval. In particular, WEFT-14B improves over Agent-World-14B by 6.41, 2.23, and 12.27 percentage points. WEFT-35B-A3B further extends these gains to more challenging long-horizon workflow benchmarks, including Toolathlon-Verified and AutomationBench.
发表机构
- East China Normal University(华东师范大学)
- Fudan University(复旦大学)
- Shanghai Innovation Institute(上海创新研究院)
- Renmin University of China(中国人民大学)
- Shanghai Qiji Zhifeng Co., Ltd(上海齐治智峰有限公司)
机构由 AI 辅助整理,请以论文原文为准。