SkillSeek:在市场规模的智能体技能检索中重新审视
SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale
- Stevens Institute of Technology(史蒂文斯理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大规模智能体技能检索中LLM介导循环的高成本问题,提出基于标准IR配方的两阶段检索器SkillSeek,在SkillsBench上以极低开销达到同等性能,成为强默认选择。
AI中文摘要:
Anthropic 的 Agent Skills 将 LLM 智能体的可复用程序性知识打包到 this http URL 目录中,开源聚合已增长超过 230,000 个技能,使得选择而非创作成为瓶颈。文献中的现有答案将选择外包给智能体本身:一个由 LLM 介导的检索循环,在智能体的决策循环内重写查询并细化候选,每个任务都需支付 LLM 令牌费用。我们提出 SkillSeek,一个基于标准 IR 配方构建的开源两阶段技能检索器(BGE-base 双编码器馈送小型交叉编码器,通过 MCP 暴露)。在 89 任务 SkillsBench 基准上的 $4 \ imes 11$ 池、骨干和方法网格中,SkillSeek 在几乎零额外成本下达到与 Liu 等人的 LLM 介导循环的观察平价:仅 bm25 在四种设置中的三种上记录了等于或高于其细化循环的通过率,小型交叉编码器在第四种上覆盖了剩余差异。第一阶段召回上限解释了该模式,每试验总支出从 51.30 美元降至 27.54 美元(在无技能基线的五十美分以内)。在我们测试的 SkillsBench 任务和 OpenHands 框架下,这将标准 IR 配方定位为智能体技能检索的强默认选择,而 LLM 介导的替代方案自然适用于确定性方法不足的情况。
英文摘要:
Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a $4 \times 11$ grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.