arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillReason:面向隐式用户请求的推理增强型智能体技能检索

SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests

Donghong Jiang, Endian Lin, Luoping Cui, Hanqing Liu, Mingjie Liu, Fan Yang, Hong Wang, Zhao Yang, Chuang Zhu

arXiv 2608.08640首次发表:更新:

AI 中文总结

针对隐式用户请求的技能检索难题,提出两阶段框架SkillReason,结合推理增强训练,在SkillReason-Bench等三个基准上取得最优性能,缩小任务目标与技能能力的语义差距。

AI 中文摘要

大型语言模型智能体越来越依赖可复用技能来扩展其超出参数知识的能力。然而,从大规模技能库中检索合适的技能仍具挑战性,因为实际用户请求往往简洁且不明确,仅说明任务目标,而将所需能力和执行步骤隐去。现有基准对这类请求的覆盖有限。为解决这一差距,我们推出SkillReason-Bench,这是一个大规模跨领域基准,包含3729个查询和61228个技能的检索语料库,涵盖9个领域。我们进一步提出SkillReason,这是一个两阶段框架,使用思维链(Chain-of-Thought, CoT)推理作为技能检索的训练时监督。在第一阶段,由更强的教师模型生成的能力推理轨迹通过对比学习、检索分布对齐和语言建模提供显式监督,鼓励检索器在其查询表示中内化能力推理。在第二阶段,检索引导的GRPO目标鼓励模型探索更适合自身能力且对检索更有效的推理轨迹。在推理时,SkillReason直接对原始查询进行编码,无需自回归CoT生成,保留仅查询检索的高效性。在SkillReason-Bench、SkillRet和SRA-Bench上的大量实验表明,SkillReason在所有三个基准上均实现了最先进的性能,证明推理增强型训练更好地弥合了高级任务目标与技能能力之间的语义差距。

英文摘要

Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are often concise and underspecified, stating only the task goal while leaving the required capabilities and execu- tion steps implicit. Existing benchmarks provide limited cov- erage of such requests. To address this gap, we introduce SkillReason-Bench, a large-scale cross-domain benchmark containing 3,729 queries and a retrieval corpus of 61,228 skills spanning nine domains. We further propose SkillRea- son, a two-stage framework that uses chain-of-thought rea- soning as training-time supervision for skill retrieval. In Stage I, capability reasoning traces generated by a stronger teacher provide explicit supervision through contrastive learning, re- trieval distribution alignment, and language modeling, en- couraging the retriever to internalize capability reasoning in its query representation. In Stage II, a retrieval-guided GRPO objective encourages the model to explore reasoning trajecto- ries better suited to its own capabilities and more effective for retrieval. At inference, SkillReason directly encodes the orig- inal query without autoregressive CoT generation, preserv- ing efficient query-only retrieval. Extensive experiments on SkillReason-Bench, SkillRet, and SRA-Bench show that Skill- Reason achieves state-of-the-art performance across all three benchmarks, demonstrating that reasoning-enhanced training better bridges the semantic gap between high-level task goals and skill capabilities.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑