arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02287cs.AI

SKT:通过经过验证的合成数据生成实现规模化的技能使用训练

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

Zelin Tan, Yiqun Zhang, Hao Li, Zhiyao Cui, Hejia Geng, Shao Zhang, Hangfan Zhang, Yang Chen, Xiaosong Wang, Lilong Wang, Zhenfei Yin, Shuyue Hu, Chen Zhang, Lei Bai

AI总结:

该研究提出SKT这一经过验证的规模化技能使用训练合成数据生成流程,构建SkillEval基准,实验证实其生成轨迹的监督微调可提升多模型技能使用性能,验证了方法的有效性与可扩展性。

AI中文摘要:

智能体技能已成为为语言模型智能体配备可复用程序知识的重要机制。然而,仅提供技能并不能保证当前模型能有效识别、应用与协调这些技能。为提升技能使用能力,我们提出SKT,这是一种经过验证的数据合成流程,可从大量智能体技能集合中构建基于技能的任务与可执行轨迹。SKT会选择合适的单技能与多技能配置,通过基于规则和智能体的验证结合反馈引导的修复来合成任务,且仅保留充分利用每项所需技能的成功轨迹。利用2000个公开技能,SKT生成了4000个任务包与27164条经过验证的轨迹。基于同一流程及不相交的测试池,我们进一步构建了SkillEval,这是一个用于评估技能使用的预留可执行基准。在不同模型、基准与智能体框架上开展的实验表明,对SKT生成的轨迹进行监督微调可持续提升技能使用性能。验证消融研究、跨框架评估与规模化实验进一步证明,这些提升依赖于高质量的监督,可扩展至单个智能体接口之外,且会随技能覆盖范围扩大而增加。综上,这些结果确立了经过验证的数据合成是一种有效且可规模化的技能使用训练方法。

英文摘要:

Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that current models can effectively identify, apply, and coordinate them. To improve skill-use capabilities, we introduce SKT, a verified data synthesis pipeline that constructs skill-grounded tasks and executable trajectories from large collections of agent skills. SKT selects suitable single-skill and multi-skill configurations, synthesizes tasks through rule-based and agent-based verification with feedback-guided repair, and retains only successful trajectories that substantially use every required skill. Using 2,000 public skills, SKT produces 4,000 task packages and 27,164 verified trajectories. Based on the same pipeline and a disjoint test pool, we further construct SkillEval, a held-out executable benchmark for evaluating skill use. Experiments across diverse models, benchmarks, and agent harnesses show that supervised fine-tuning on SKT-generated trajectories consistently improves skill-use performance. Verification ablations, cross-harness evaluation, and scaling experiments further demonstrate that these gains depend on high-quality supervision, extend beyond a single agent interface, and increase with broader skill coverage. Together, these results establish verified data synthesis as an effective and scalable approach for skill-use training.

补充信息

↑