arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37105cs.LGcs.AIcs.CL

VACE:智能体模型与框架的验证门控交替协同进化

VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses

  • ICT AI Competence Center, Huawei Technologies Co., Ltd.(华为技术有限公司ICT AI能力中心)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Jiexing Qi, Yu He, Jun Liu, Qichen Huang, Shaohua Hu, Zhan Dang, Guohua Chen, Rui Yang, Wen Jiang, Yang Liu, Tao Lyu, Fangming Li

AI总结:

VACE通过交替进行智能体强化学习与轨迹驱动的框架细化,并采用验证门控机制,在Qwen3.5-9B上显著提升了OfficeQA和AutomationBench的性能,优于仅权重更新和未加门控的交替方法。

AI中文摘要:

语言模型智能体可以通过更新其模型权重或改进引导任务执行的框架(harness)来得到提升。这两个组成部分是相互耦合的:权重更新会改变模型使用框架的方式,而框架更新则会改变用于训练的数据轨迹。我们提出了VACE(验证门控交替协同进化,Validation-Gated Alternating CoEvolution),该方法将智能体强化学习与轨迹驱动的框架细化交替进行。在每个强化学习阶段之后,VACE会重用收集到的轨迹来提出一个框架修订方案,并在保持更新后的模型固定不变的情况下,评估现有方案和候选方案。仅当候选方案能提升验证性能时,它才会指导后续的训练。使用Qwen3.5-9B模型,VACE在OfficeQA上达到了45.26%的测试准确率,在AutomationBench上达到了75.19%的平均部分得分,分别比仅进行权重更新的强化学习高出6.43和9.09个百分点,比未加门控的交替方法高出4.59和6.95个百分点。在44个框架提案中,有17个在更新后的检查点上降低了验证性能,并在后续强化学习训练之前被拒绝,这凸显了验证门控的重要性。

英文摘要:

Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoEvolution, which alternates agentic reinforcement learning with trajectory-driven harness refinement. After each RL stage, VACE reuses the collected trajectories to propose a harness revision and evaluates the incumbent and candidate with the updated model held fixed. The candidate guides subsequent training only if it improves validation performance. With Qwen3.5-9B, VACE achieves 45.26% test accuracy on OfficeQA and a mean partial-credit score of 75.19% on AutomationBench, exceeding weight-only RL by 6.43 and 9.09 percentage points and ungated alternation by 4.59 and 6.95 points, respectively. Across 44 harness proposals, 17 reduce validation performance at the updated checkpoint and are rejected before subsequent RL training, highlighting the importance of validation gating.

↑