arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37539cs.AI

SkillGym:通过自动可验证环境生成训练技能使用智能体

SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

Renxi Wang, Mingshan Hee, Fajri Koto, Timothy Baldwin, Haonan Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对技能使用训练数据不足问题,提出SkillGym自动流水线,构建可验证环境并收集轨迹,微调后显著提升多规模LLM的技能调用率与基准性能。

中文摘要 AI 辅助

技能赋予LLM智能体专业知识和指导,以完成长周期和复杂任务。尽管技能已在近期的智能体范式和工具中被广泛采用,但如何合成可靠的训练数据以及如何训练智能体使用技能仍未被充分探索。在本工作中,我们提出SkillGym,一个自动流水线,用于构建可验证环境、收集轨迹并训练技能使用智能体。SkillGym首先从互联网抓取大量技能,然后保留那些工作流可在离线条件下可复现运行的技能。采用构建者-审查者流水线来构造难度可控的任务,涵盖四种任务类型,每种任务都有参考解决方案和可执行验证器。利用该流水线,我们构建了6.8k个环境,并收集了19k条经验证的成功轨迹用于监督微调。在这些轨迹上进行微调,提升了不同家族和规模的LLM(从2B到122B参数)在四个技能使用基准上的表现;我们的Qwen3.5-9B SFT模型在其中两个基准上优于未训练的397B模型。进一步分析表明,训练教会了智能体调用技能,将读取相关技能的比例从28%提升至96%,并且这些收益在不同推理结构上保持,扩展到构成训练数据少数的任务类型以及训练中未见的技能。

英文摘要

Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to construct difficulty-controlled tasks, spanning four task types, each with a reference solution and an executable verifier. With this pipeline, we build 6.8k environments and collect 19k verified successful trajectories for supervised finetuning. Finetuning on these trajectories improves LLMs of different families and sizes, from 2B to 122B parameters across four skill-use benchmarks; Our Qwen3.5-9B SFT model outperforms the 397B untrained model on two of them. Further analysis shows that training teaches agents to invoke skills, raising the rate of reading the relevant skill from 28% to 96%, and that the gains hold across reasoning structures, extending to task types that form a minority of the training data and to skills held out from training

发表机构

  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

↑