arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10538cs.AI

SKILLER:面向小型语言模型可复用技能提取的语言级强化学习框架

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

Chenhao Dang, Siyuan Xiong, Conghui He, Weijia Li

中文总结 AI 辅助

研究针对小型语言模型技能生成成本高的问题,提出SKILLER框架,经实验在多基准上优于现有方法,性能接近闭源模型,可大幅降低智能体技能部署成本。

中文摘要 AI 辅助

智能体技能是封装过程性知识与领域专长的标准化格式,在智能体管控系统中作为持续约束语言模型行为空间以实现可重复、高质量任务执行的核心机制。但由于强大的闭源模型推理成本高昂,当前流行的智能体管控系统(如Codex和OpenClaw)在部署这些技能完成实际任务时仍成本过高。可在消费级GPU上部署的开源模型能力快速提升,为利用基于技能的行为约束大幅降低成本提供了契机。然而,针对这类紧凑模型自动生成有效技能仍是重大实践挑战。为解决该问题,我们提出SKILLER——一种自然语言驱动的强化学习框架,专为小型模型自动生成特定执行器的技能,该框架采用强大模型作为演员与评论员,将小型模型智能体系统视为环境,且完全通过自然语言传播所有强化学习信号。在五个相关基准上使用Qwen3.5-9B和Qwen3.5-4B开展的大量实验评估表明,SKILLER的性能优于三种开源和一种闭源技能生成或演化方法,9B模型的绝对提升幅度为4.3至20.4个百分点,4B模型为1.8至13.3个百分点,且在SkillsBench的单技能任务上,其性能显著匹配强大闭源模型的表现。该项目可在指定URL获取。

英文摘要

Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high-quality task execution. However, because strong closed-source models entail high inference costs, current popular agent harnesses, such as Codex and OpenClaw, remain prohibitively expensive when deploying these skills to accomplish real-world tasks. The rapid capability enhancement of open-source models deployable on consumer-grade GPUs presents a compelling opportunity to drastically reduce these costs by leveraging skill-based behavioral constraints. Nevertheless, automatically generating effective skills tailored specifically for such compact models remains a significant practical challenge. To address this, we propose SKILLER, a natural-language-driven reinforcement learning framework designed to automatically generate executor-specific skills for small models, which employs a strong model as the actor and critic, treats the small-model agent system as the environment, and propagates all reinforcement learning signals entirely via natural language. Extensive experimental evaluations across five relevant benchmarks using Qwen3.5-9B and Qwen3.5-4B demonstrate that SKILLER outperforms three open-source and one closed-source skill generation or evolution methods, achieving absolute gains ranging from 4.3 to 20.4 percentage points for the 9B model and 1.8 to 13.3 points for the 4B model, while remarkably matching the performance of strong closed-source models on single-skill tasks in SkillsBench. The project is available at https://github.com/DANG-ai/SKILLER.

补充信息

↑