arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22695cs.CLcs.AIcs.IR

Enrich-Retrieve-Rank:将能力发现扩展至上下文路由之外

Enrich-Retrieve-Rank: Scaling Capability Discovery Beyond In-Context Routing

  • Amazon AGI(亚马逊AGI)

机构由 AI 辅助整理,请以论文原文为准。

Nazib Sorathiya, Daniel Zhang, Bardiya Akhbari

AI总结:

该研究提出Enrich-Retrieve-Rank流程,将能力发现从上下文路由扩展,经实验验证其在大规模MATS组件场景下,比Full-Ctx、Search&Pick基线性能更优且成本更低,已作为多智能体平台的默认能力发现层投入生产。

AI中文摘要:

智能体生态系统如今包含数千种MATS组件(模型、智能体、工具和技能),但它们的发现仍依赖于上下文路由。这些系统会读取注册表(包含名称、提示或描述,具体取决于上下文预算),选择一个候选组件,调用它,并在失败时重试。这种模式会随着规模扩大而性能下降,且注册表正在快速增长。我们将能力发现重新定义为对注册表的搜索,方法是定义一个离线富集步骤,该步骤将稀疏元数据转换为可搜索的配置文件,以及一个在线检索-然后-排序流程,该流程会返回一个排名靠前的短名单,无需在线调用任何候选组件。我们展示了当能力数量N从10增加到7278时,上下文路由的Top-1准确率(Match@1)大幅下降(从0.85降至0.12),而检索-然后-排序的性能下降更为平缓(从0.81降至0.39),这是因为一旦检索到正确的能力,其重排器仍能以0.70至0.87的概率将其排在首位。在Nova Micro扫描中,交叉点约为N=500。我们与两个上下文基线进行了比较:Full-Ctx将整个注册表放入提示中并要求大语言模型进行选择;Search&Pick为大语言模型提供一个搜索工具,在选择前缩小候选范围。在全规模下,该流程在Match@1上比Search&Pick领先6.5个百分点(pp),成本仅为其一半左右,相比Full-Ctx则降低了70倍成本。我们在智能体、工具和技能注册表中使用相同的固定配置(相同的富集、检索器和评分器权重),该流程已作为大型多智能体平台的默认能力发现层投入生产运行。

英文摘要:

Agent ecosystems now include thousands of MATS components (Models, Agents, Tools, and Skills), yet their discovery still relies on in-context routing. These systems read a registry (names, hints, or descriptions, as context budget permits), pick a candidate, invoke it, and retry on failure. This pattern degrades with scale, and registries are growing fast. We recast capability discovery as search over a registry by defining an offline enrichment step that turns sparse metadata into searchable profiles, and an online retrieve-then-rank pipeline that returns a ranked shortlist without invoking any candidates online. We show that from N=10 to 7,278 capabilities, in-context routing's top-1 accuracy (Match@1) collapses (0.85 to 0.12), while retrieve-then-rank degrades more gently (0.81 to 0.39) because its reranker still ranks the right capability first 0.70-0.87 of the time once retrieval finds it. In the Nova Micro sweep, the crossover is around N=500. We compare against two in-context baselines. Full-Ctx puts the whole registry in the prompt and asks the LLM to pick. Search&Pick gives the LLM a search tool to narrow candidates before it picks. At full scale the pipeline leads Search&Pick by 6.5 percentage points (pp) on Match@1 at about half the cost. It reduces cost 70x versus Full-Ctx. We use a fixed configuration (same enrichment, retriever, and scorer weights) across agent, tool, and skill registries. The pipeline runs in production as the default capability-discovery layer of a large-scale multi-agent platform.

补充信息

↑