arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10426cs.CL

CoTrace:基于模型与运行框架协同进化的终端智能体训练数据配方

CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

Jixuan Chen, Jiaxin Zhang, Qinyuan Ye, Yada Pruksachatkun, Haoxiang Zhang, Jingming Zhuo, Yifan Zhang, Yutong Dai, Juntao Tan, Xiangyu Peng, Silvio Savarese, Ze… 展开作者

Jixuan Chen, Jiaxin Zhang, Qinyuan Ye, Yada Pruksachatkun, Haoxiang Zhang, Jingming Zhuo, Yifan Zhang, Yutong Dai, Juntao Tan, Xiangyu Peng, Silvio Savarese, Zeyuan Chen, Lianhui Qin, Chien-Sheng Wu

首次发表
浏览论文内容

中文总结 AI 辅助

CoTrace提出一种框架感知的数据配方,通过轨迹路由、来源匹配和课程刷新优化终端智能体训练,在Tmax上使Qwen3.5-9B从78提升至88个任务,并验证了框架兼容性对分布外迁移的关键作用。

中文摘要 AI 辅助

终端智能体的能力同时取决于模型权重和运行时框架(runtime harness),后者负责格式化提示、绑定工具和处理错误恢复。现有的框架-模型协同进化方法虽然能同时改进这两个组件,但往往将在框架搜索过程中产生的轨迹视为无差别的回放缓冲区。这种做法忽略了轨迹对模型训练的价值取决于其生成时所处的框架这一事实。为了系统地分析这一接口,我们建立了一个交替协同进化框架,通过组件级提升决策将框架搜索与策略训练解耦。在此框架内,我们提出了CoTrace,一种框架感知的数据配方,明确管理轨迹路由、来源匹配和课程刷新。在CoTrace下,反复出现的执行失败指导框架合成,而策略训练严格基于与所采用运行时匹配的已验证回放(用于监督微调SFT)或新的在线交互(用于强化学习RL)进行条件化。在Tmax提升划分上,CoTrace在监督微调下将Qwen3.5-9B从解决78个任务提升至88个,而在线强化变体达到90个。具体而言,一个紧凑的框架匹配语料库以显著低于跨兄弟框架汇集的大型语料库的计算成本,产生了稳定的模型性能提升。此外,在Terminal-Bench 2.1和SWE-bench Lite上的评估表明,分布外迁移从根本上依赖于框架兼容性,保持训练和评估运行时之间的一致性可防止在外部脚手架下观察到的程序性执行故障。

英文摘要

Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This practice overlooks that a trajectory's value for model training depends on the harness under which it was generated. To systematically analyze this interface, we establish an alternating co-evolution framework that decouples harness search and policy training through component-wise promotion decisions. Within this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh. Under CoTrace, recurring execution failures guide harness synthesis, while policy training is strictly conditioned on verified rollouts matched to the adopted runtime for supervised fine-tuning (SFT) or fresh online interactions for reinforcement learning (RL). On the Tmax promotion split, CoTrace advances Qwen3.5-9B from 78 to 88 solved tasks under supervised fine-tuning while an online reinforcement variant reaches 90. Specifically, a compact harness-matched corpus produces steady model gains at substantially lower compute than much larger corpora pooled across sibling harnesses. Furthermore, evaluations on Terminal-Bench 2.1 and SWE-bench Lite show that out-of-distribution transfer depends fundamentally on harness compatibility, where maintaining consistency between training and evaluation runtimes prevents procedural execution breakdowns observed under foreign scaffolds.

发表机构

  • University of California, San Diego(加利福尼亚大学圣迭戈分校)
  • Salesforce AI Research(赛富时人工智能研究院)
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑