arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

技能了解其邻居:用于技能检索的聚类对比能力页面

Skills Know Their Neighbors: Cluster-Contrastive Capability Pages for Skill Retrieval

Zifei Wang, Wei Wen, Qiang Ji, Ruizhi Qiao

arXiv 2608.04482首次发表:更新:

AI 中文总结

针对大语言模型智能体技能检索中仅改进检索器无法消除的文档导致的误差,提出聚类对比的能力页面,在SRA-Bench等数据集上显著提升了召回率与任务成功率。

AI 中文摘要

随着技能库的扩大,大语言模型智能体必须从候选技能中检索可复用的技能,这些候选技能往往具有相同的主题和词汇,但实现不同的能力。检索不仅受评分器限制,还受待评分文本限制:文档可能描述技能的作用,却未说明哪些相似请求应路由到其他地方。我们将技能的能力形式化为其\textit{可执行区域},即它能解决的查询集合,并将其文档视为该区域的有损观测。这一观点揭示了检索误差中由文档导致的部分,仅通过改进检索器无法消除。因此,我们提出\textit{能力页面},这是一种聚类对比的技能表示,包含正触发项$\tau_{pos}$、负边界$\tau_{neg}$和判别主体$B$。离线编译器会比较相邻技能以生成这些字段。推理时,索引使用$\tau_{pos}$和$B$进行候选召回,而路由器使用$\tau_{neg}$拒绝易混淆的替代项。在包含26262个技能和来自6个数据集的5400个问题的SRA-Bench上,能力页面提升了所有5种测试检索器的Recall@10,平均提升2.94个百分点;向候选卡片添加$\tau_{neg}$后,在4个执行器和6个数据集上的端到端任务成功率平均提升3.62个百分点。在中文SSL-SkillDiscovery上的迁移评估中,使用相同编码器在各条件下达到73.07%的MRR@50。能力页面无需修改在线模型,仅通过重写离线技能库即可提升路由效果。

英文摘要

As skill libraries grow, large language model agents must retrieve reusable skills from candidates that often share the same topic and vocabulary but implement different capabilities. Retrieval is limited not only by the scorer but also by the text being scored: a document may describe what a skill does without stating which similar requests should be routed elsewhere. We formalize a skill's capability as its \emph{executable region}, the set of queries it can solve, and view its document as a lossy observation of that region. This view exposes a document-imposed component of retrieval error that cannot be removed by improving the retriever alone. We therefore propose \emph{Capability Pages}, cluster-contrastive skill representations containing a positive trigger $\Tpos$, a negative boundary $\Tneg$, and a discriminative body $B$. An offline compiler compares neighboring skills to write these fields. At inference time, the index uses $\Tpos$ and $B$ for candidate recall, while the router uses $\Tneg$ to reject confusable alternatives. On SRA-Bench, which contains 26{,}262 skills and 5{,}400 questions from six datasets, Capability Pages improve Recall@10 for all five tested retrievers, with a mean gain of $2.94$ points. Adding $\Tneg$ to candidate cards improves end-to-end task success by $3.62$ points on average across four executors and six datasets. A transfer evaluation on Chinese SSL-SkillDiscovery reaches $73.07\%$ MRR@50 using the same encoder across conditions. Capability Pages require no modification to the online models; they improve routing by rewriting the offline skill library.

Comments16 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑