arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

元任务:将终端任务综合转化为可扩展智能体训练的终端任务

Meta-Task: Turning Terminal Task Synthesis into a Terminal Task for Scalable Agent Training

Zhihong Pan, Jiyuan He, Kai Zhang, Yupeng Han, Ze Liu, Yuze Zhao, Yongcong Ye, Zhaohua Yang

arXiv 2607.27929首次发表:更新:

发表机构

State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China; Meituan(中国科学技术大学; 美团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Meta-Task框架将终端任务综合转化为Terminal-Bench格式任务,通过多阶段机制等提升任务质量,生成3221条轨迹微调Qwen3系列模型,效果优于同期方法且训练数据更少。

AI 中文摘要

大规模训练终端智能体需要多样化、可验证的终端任务及高质量交互轨迹,但获取此类数据仍是重大挑战。现有综合方法存在两大核心局限:(1)任务生成与实际执行脱节导致可靠性弱;(2)依赖现有仓库导致多样性和可扩展性有限。我们提出Meta-Task框架,将终端任务综合本身重新定义为Terminal-Bench格式的任务:智能体在真实容器环境中运行,迭代生成、执行并验证任务,使生成的组件在生成循环内接受内部一致性和可执行性检查。在此基础上,我们沿多个维度解耦目标任务需求,引入多阶段机制,在生成实际任务前动态设计新颖的任务规格,并纳入可选的外部材料支持以提升多样性和真实性。我们还应用LLM-as-Judge过滤确保最终训练数据的质量。在Terminal-Bench 2.0上的实验显示,仅用3221条Meta-Task生成的轨迹进行微调,Qwen3-14B和Qwen3-32B分别达到22.5%和31.8%的Avg Pass@1,在训练数据显著更少的情况下优于同期方法。

英文摘要

Training terminal agents at scale requires diverse, verifiable terminal tasks and high-quality interaction trajectories, yet acquiring such data remains a significant challenge. Existing synthesis methods face two key limitations: (1) weak reliability caused by the disconnect between task generation and real execution, and (2) limited diversity and scalability due to dependence on existing repositories. We propose Meta-Task, a framework that redefines terminal task synthesis as a Terminal-Bench-format task itself: an agent operates within a real container environment to iteratively generate, execute, and verify tasks, so that synthesized components are checked for internal consistency and executability within the generation loop itself. Building upon this, we decouple the target task requirements along multiple dimensions, introduce a multi-phase mechanism that dynamically designs novel task specifications before producing the actual tasks, and incorporate optional external material support to enhance diversity and realism. We additionally apply LLM-as-Judge filtering to ensure the quality of the final training data. Experiments on Terminal-Bench 2.0 show that fine-tuning on only 3,221 Meta-Task synthesized trajectories achieves 22.5% and 31.8% Avg Pass@1 for Qwen3-14B and Qwen3-32B respectively, outperforming concurrent approaches with significantly less training data.

Comments17 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑