arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05466cs.AIcs.LG

长视界终端任务的递归合成

Recursive Synthesis for Long-Horizon Terminal Tasks

Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出RST递归合成框架,低成本规模化生成3.7万余个长视界终端任务,经微调或PPO优化后,多款Qwen模型在多个终端基准任务上性能显著提升,且递归过程无明显上限。

中文摘要 AI 辅助

为终端智能体生成高质量的长视界训练数据成本高昂,每个任务需保证指令、环境、参考解和验证器相互一致,单任务成本常达数百至数千美元,人工编写无法规模化,直接用大语言模型(LLM)生成常破坏这些依赖关系。本文提出递归合成终端任务框架(Recursive Synthetic Terminal Tasks, RST),用于规模化构建长视界终端智能体任务。RST从已验证的种子任务出发,扩展参考解,将验证器和指令重新适配新工作流,在全新沙箱中验证结果,并将已接受的任务作为种子用于后续轮次。经过15轮递归,RST生成37484个合成终端智能体任务,单任务成本约0.05美元。任务难度随轮次大幅提升:参考解中位数从67行增至374行,执行命令中位数从40条增至244条,DeepSeek-V4-Pro的pass@4指标从第1轮的90%降至第15轮的2.5%。为验证训练效用,我们在合成任务上收集拒绝采样的Qwen3.5轨迹并用于监督微调,在这些轨迹上微调可使Qwen3.5-27B和Qwen3.5-122B-A10B在Terminal-Bench 2、Terminal-Bench Hard和Long-Horizon Terminal Bench上的性能提升最高达10个百分点;同时,智能体PPO算法将Qwen3.5-27B在上述三个基准上的性能分别提升至49.44%、32.00%和22.07%,较基础模型的相对提升分别为20.0%、41.2%和21.9%。此外,15轮递归后仍未出现性能上限:合成产量和验证率保持稳定,难度持续攀升,表明该过程可远超本文报告的规模继续推进。

英文摘要

High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consistent. Human authoring does not scale, and direct generation with large language models (LLMs) often breaks these dependencies. We present Recursive Synthetic Terminal Tasks (RST), a recursive verified synthesis framework for constructing long-horizon terminal-agent tasks at scale. Starting from verified seed tasks, RST extends the reference solution, realigns the verifier and instruction to the new workflow, validates the result in a fresh sandbox, and reuses accepted tasks as seeds for subsequent rounds. Across fifteen recursive rounds, RST produces 37,484 synthesized terminal-agent tasks at roughly $0.05 per task. Task difficulty increases substantially over rounds: the median reference solution grows from 67 to 374 lines, the median number of executed commands grows from 40 to 244, and DeepSeek-V4-Pro pass@4 drops from 90% at R1 to 2.5% at R15. To demonstrate training utility, we collect rejection-sampled Qwen3.5 trajectories on the synthesized tasks and use them for supervised fine-tuning. Fine-tuning on these trajectories improves Qwen3.5-27B and Qwen3.5-122B-A10B by up to 10 points on Terminal-Bench 2, Terminal-Bench Hard, and Long-Horizon Terminal Bench, while agentic PPO lifts Qwen3.5-27B to 49.44%, 32.00%, and 22.07% on the three benchmarks, corresponding to relative gains of 20.0%, 41.2%, and 21.9% over the base model. Moreover, after 15 rounds, the recursion shows no ceiling: synthesis yield and validation rates remain stable as difficulty keeps climbing, indicating that the process can continue well beyond the scale reported here.

发表机构

  • Tencent HY LLM Frontier(腾讯HY大模型前沿团队)
  • University of Georgia(佐治亚大学)
  • University of Maryland, College Park(马里兰大学帕克分校)
  • University of Pennsylvania(宾夕法尼亚大学)
  • University of Minnesota, Twin Cities(明尼苏达大学双城分校)
  • Indiana University(印第安纳大学)
  • National University of Singapore(新加坡国立大学)
  • Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑