arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPT:将技能作为智能体语言模型的预训练数据

SPT: Skills as Pre-Training Data for Agentic Language Models

Yufei Sun, Yudong Li, Yiming Cheng

arXiv 2608.26563首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; Tsinghua University(北京邮电大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出技能预训练(SPT)方法,将公开技能包作为训练数据,结合参考感知组装策略,可提升智能体语言模型的工具使用性能,同时保留通用性能,验证了技能包作为预训练数据源的价值。

AI 中文摘要

智能体(使用工具的)语言模型主要在后期训练阶段基于工具调用轨迹和智能体轨迹进行训练,这些数据提供了直接的行为监督,但生成它们需要任务环境、执行和验证,因此难以覆盖广泛的工具和任务,成本较高。公开技能提供了另一种训练数据源:它们编码了可复用的工具语义和工作流,但通常仅用作推理时的上下文。我们提出技能预训练(Skill Pre-Training,SPT),这是一种中期训练方法,将因果语言建模应用于SkillCorpus(一个公开的多文件技能包集合),可选择性地与通用数据混合。为保留每个包内文件之间的关系,我们还提出Reference Insert,一种参考感知的组装策略,将支持文件放置在其在主指令中的提及附近。在多个模型规模和后期训练方案上的实验表明,与基于通用数据或轨迹数据的中期训练相比,SPT能持续提升智能体性能,同时基本保留通用性能。数据混合实验显示,将技能数据与通用退火语料库结合还能带来额外收益。这些结果表明,技能包是预训练智能体语言模型的宝贵数据源。

英文摘要

Agentic (tool-using) language models are mainly trained on tool-call traces and agent trajectories during post-training. These data provide direct behavioral supervision, but producing them requires task environments, execution, and verification, making broad tool and task coverage expensive. Publicly available skills offer another source of training data: they encode reusable tool semantics and workflows but are typically used only as inference-time context. We introduce Skill Pre-Training (SPT), a mid-training method that applies causal language modeling to SkillCorpus, a collection of public multi-file skill packages, optionally mixed with general data. To preserve relations among files within each package, we also introduce Reference Insert, a reference-aware assembly strategy that places supporting files near their mentions in the primary instruction. Experiments across multiple model scales and post-training recipes show that SPT consistently improves agentic performance over mid-training on general or trajectory data, while largely preserving general performance. Data mixture experiments show additional benefits from combining skill data with general annealing corpora. These results indicate that skill packages are a valuable data source for pre-training agentic language models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑