arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大型技能库的智能体检索方法的比较研究

Comparative Approaches to Agent Retrieval over Large Skill Libraries

Indivara Kolluru, Nathan Sportsman

arXiv 2608.06196首次发表:更新:

发表机构

Praetorian(Praetorian)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对比了混合排名器与类型化知识图谱在690个技能库的117个真实查询上的检索效果,发现添加结构无法提升强排名器的检索性能,明确了结构优化的适用条件。

AI 中文摘要

拥有大型技能库的智能体必须决定加载哪些技能以及加载顺序,将整个库加载到上下文的成本很高,且无法为自主排序提供结构。我们针对包含690个技能的语料库,研究了两种解决该问题的系统:一种是混合排名器,结合了词汇检索和密集嵌入检索,用于稀疏的按需加载;另一种是类型化知识图谱,对先决条件、数据流和排序等工作流关系进行编码。在117个真实、非重复的查询上,混合排名器在73.5%±8.0的情况下能在前5个结果中检索到正确技能,约四分之一的查询无法被覆盖。当按设计意图使用该图谱(在匹配的令牌预算下,用图谱邻居替代额外的排名结果)时,其性能显著更差(下降11.2个百分点,p=0.0007);其大语言模型(LLM)生成的边层相比从本地嵌入通道免费获取的邻居没有任何增益,且排名器遗漏的73%的查询根本无法通过该图谱访问。我们将此归因于预过滤拓扑边界:由于图谱的候选边来自排名器已搜索的同一嵌入邻域,98.6%的类型化边连接着排名器已共同呈现的技能。该图谱可丰富关系语义,但无法扩展检索范围。我们进一步表明,基于作者编写的查询进行评估会将命中@5夸大多达44个百分点,这会完全掩盖上述结果。我们的贡献是对为何添加结构无法提升强大排名器的检索性能给出了机制性解释,并确定了在检索中加入结构相互依存关系达到最优的条件。

英文摘要

Agents backed by large skill libraries must decide which skills to load and in what order. Loading the entire library into context is expensive and provides no structure for autonomous sequencing. We study two systems for this problem over a corpus of 690 skills: a hybrid ranker combining lexical and dense-embedding retrieval for sparse, on-demand loading, and a typed knowledge graph encoding workflow relations such as prerequisites, data flow, and ordering. On a set of 117 realistic, non-echoing queries, the hybrid ranker retrieves the correct skill within the top five in 73.5% +/- 8.0 of cases, leaving roughly a quarter of queries unserved. When used as the design intended (substituting graph neighbours for additional ranked results at matched token budget), the graph is significantly worse (-11.2 points, p = 0.0007). Its LLM-generated edge layer adds nothing over neighbours obtained free from a local embedding pass, and 73% of the queries the ranker misses are not reachable through the graph at all. We attribute this to a pre-filter topology bound. Because the graph's candidate edges are drawn from the same embedding neighbourhood the ranker already searches, 98.6% of typed edges connect skills the ranker had already surfaced together. The graph can enrich relation semantics but cannot extend retrieval reach. We further show that evaluating on author-written queries overstates hit@5 by up to 44 points, which would have hidden these results entirely. Our contribution is a mechanistic account of why added structure does not improve retrieval over a strong ranker, and identify the conditions under which adding structural interdependence into the retrieval is optimal.

Comments9 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑