arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16071cs.CLcs.IR

Skill2Query:利用技能结构生成智能体技能检索的伪查询

Skill2Query: Exploiting Skill Structure to Generate Pseudo-Queries for Agent Skill Retrieval

Lihui Ding, Zihan Guo, Bingwei Lu, Chenyu Zhou, Yuanjian Zhou, Weinan Zhang, Jianghao Lin, Dongdong Ge

首次发表
浏览论文内容

中文总结 AI 辅助

Skill2Query是利用技能知识图谱生成伪查询的框架,可提升多类检索性能,生成的训练数据表现最优,还能提高智能体任务成功率。

中文摘要 AI 辅助

伪查询生成可缓解智能体技能检索的监督瓶颈,但现有文档级方法通常未明确利用能力、参数和使用示例之间丰富的内部关系。因此,生成的查询可能在主题上与技能相关,却缺乏能力基础和参数一致性,这引发了一个问题:明确利用技能文档的内部结构是否能产生更有效的检索信号?为此,我们提出Skill2Query框架,该框架首先将技能文档解析为技能知识图谱,随后通过风格模仿、查询模板生成和参数填充三个阶段生成伪查询。生成的查询可用于离线索引增强、在线查询扩展和检索器训练。我们使用四个基准(TheoremQA、LogicBench、ToolQA和CHAMP),结合跨多个下游应用的大规模技能候选池(包括技能检索、检索器训练和端到端智能体执行)对Skill2Query进行评估。我们生成了涵盖不同领域的近3万个技能,共70万个类别多样的伪查询。Skill2Query在稀疏检索、密集检索和技能路由检索中均实现了性能提升,在各类检索设置中平均Recall@1提升6.70个百分点。Skill2Query生成的训练数据在评估的所有生成基线中取得了最佳的Recall@1和nDCG@1。对多个大语言模型(LLM)后端的进一步评估表明,技能检索性能的提升可转化为更高的智能体任务成功率。代码和资源可在该https URL获取。

英文摘要

Pseudo-query generation can alleviate the supervision bottleneck for agent skill retrieval, but existing document-level approaches typically leave the rich internal relations among capabilities, parameters, and usage examples implicit. As a result, generated queries may be topically relevant to a skill while lacking capability grounding and parameter consistency, raising the question of whether explicitly exploiting a skill document's internal structure can produce more effective retrieval signals. We therefore propose Skill2Query, a framework that first parses a skill document into a Skill Knowledge Graph and then generates pseudo-queries through a three-stage process including style mimicking, query template generation, and parameter filling. The generated queries can be used for offline index augmentation, online query expansion, and retriever training. Four benchmarks (TheoremQA, LogicBench, ToolQA, and CHAMP) are used to evaluate Skill2Query with large-scale skill candidate pools across multiple downstream applications, including skill retrieval, retriever training, and end-to-end agent execution. Using nearly 30K skills across diverse domains, we generate 700K category-diverse pseudo-queries. Skill2Query consistently improves sparse, dense, and skill-routing retrieval, with an average Recall@1 gain of 6.70 percentage points across retrieval settings. Skill2Query-generated training data also achieves the best Recall@1 and nDCG@1 among the evaluated generation baselines. Further evaluations with multiple LLM backends demonstrate that improved skill retrieval translates into higher agent task success rates. Code and resources are available at https://github.com/MatZaharia/Skill2Query.

发表机构

  • Fudan University(复旦大学)
  • Sun Yat-sen University(中山大学)
  • Shanghai Innovation Institute(上海创新研究院)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑