AI 中文总结
该研究针对现有技能构建受限于模型已有知识的问题,提出Search2Skill框架,通过基于rubric的强化学习实现外部证据蒸馏为可复用技能,在多领域实验中优于基线且技能可跨模型规模迁移。
AI 中文摘要
可复用技能封装了解决现实世界专业任务所需的程序性知识,为基于大语言模型(LLM)的智能体提供了在专家领域实现自我进化的路径。现有的自我进化技能方法从模型的参数知识或轨迹内部构建技能,因此受限于模型已有的知识。然而,专业技能背后的领域惯例和标准流程往往超出该边界,难以仅从智能体本身获取。为解决此问题,我们提出了一种新框架 Search2Skill,该框架可自动识别智能体的能力缺口,搜索外部资源以填补缺口,并将检索到的证据蒸馏为结构化、可复用的技能。具体而言,Search2Skill 通过基于 rubric(评分规则)的强化学习方案进行优化,该方案在何时搜索、如何搜索以及如何生成技能三个方面共同优化。在来自三个基准的八个专家级领域上进行的实验表明,Search2Skill 在流式评估和保留评估两种协议下,均始终优于基于搜索增强和轨迹的技能学习基线。进一步分析显示,性能提升源于技能抽象而非原始检索证据,且习得的技能可跨模型规模迁移。
英文摘要
Reusable skills, which encapsulate the procedural knowledge required to solve real-world professional tasks, offer LLM-based agents a path toward self-evolution in expert domains. Existing self-evolving skill methods construct skills internally from the model's parametric knowledge or trajectories, and are therefore bounded by what the model already knows. However, the domain conventions and standard procedures underlying professional skills often lie beyond this boundary and are hard to elicit from the agent alone. To address this issue, we therefore propose a novel framework, Search2Skill, that automatically identifies the agent's capability gaps, searches external sources to address them, and distills the retrieved evidence into structured, reusable skills. Specifically, Search2Skill is optimized by a rubric-based reinforcement learning scheme that jointly improves when to search, how to search, and how to generate skills. Experiments on eight expert-level domains from three benchmarks show that Search2Skill consistently outperforms both search-augmented and trajectory-based skill-learning baselines under both streaming and held-out evaluation protocols. Further analyses show that the gains arise from skill abstraction rather than raw retrieved evidence, and that the acquired skills transfer across model scales.