AI 中文总结
研究针对大语言模型智能体技能检索问题,提出SkillSight框架,通过语义和词汇空间校准共享背景,减少偏差,实验证明该方法能提升检索指标,在端到端评估中表现出色且速度快,实现准确高效的技能选择。
AI 中文摘要
随着大语言模型智能体可访问的技能库越来越大,正确检索技能对于可靠的能力选择和执行至关重要。现有检索器常将技能描述视为普通文档,忽略其高度规则结构。共享描述模式在许多技能中反复出现,却难以区分所需能力。研究表明共享描述背景导致密集相关分数、查询与技能文档间出现能量差距并模糊任务相关信号。基于此提出SkillSight,一个无需训练的检索框架,在语义和词汇空间校准共享背景。语义背景校准从IDF识别的通用令牌估计背景子空间,减少共享描述模式引起的相似性,词汇证据校准降低共享背景令牌权重以恢复判别性令牌级证据。在SRA - Bench和SkillBench - Supp上的实验表明,SkillSight在检索指标上有持续改进,在端到端评估中,在三个智能体模型中实现最佳整体性能,比Dense + Reranker基线快1248倍。结果表明共享描述背景是技能检索偏差的关键来源,明确校准可实现准确高效的技能选择且无需额外训练。
英文摘要
As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable capability selection and execution. Existing retrievers often treat skill contents as ordinary documents, overlooking their highly regular structure: shared descriptive patterns recur across many skills while providing little evidence for distinguishing the required capability. We show that this shared descriptive background is reflected in dense relevance scores, induces a pronounced energy gap between queries and skill documents, and obscures discriminative signals, especially for structurally similar hard negatives. Based on this observation, we propose SkillSight, a training-free retrieval framework that calibrates shared background in both semantic and lexical spaces. Semantic Background Calibration estimates a background subspace from generic tokens identified by IDF, reducing similarity induced by shared descriptive patterns, while Lexical Evidence Calibration downweights shared background tokens to recover discriminative token-level evidence. Experiments on SRA-Bench and SkillBench-Supp demonstrate consistent improvements across retrieval metrics, with SkillSight improving Recall@10 by up to 20.21 percentage points over the original dense retriever. It is up to 1,248 times faster than the Dense + Reranker baseline. In end-to-end evaluation, SkillSight achieves the best overall performance across three agent models and outperforms LLM Selection by up to 4.97 percentage points. These results identify shared descriptive background as a source of ranking interference in skill retrieval and demonstrate that calibrating it enables accurate and efficient skill selection without additional training. Our code can be found at https://github.com/xiaojinying/SkillSight.
Comments9 pages, 4 figures