发表机构
Valyu AI; University of Warwick; University College London(Valyu AI; 华威大学; 伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出用小型语言模型(SLM)结合渐进式监督微调与强化学习的方法,解决多智能体路由中仅基于意图选择的局限,在智能体-查询不匹配子集及总体NDCG@10指标上均优于仅基于意图路由的大语言模型基线,且选择延迟更低。
AI 中文摘要
专用检索智能体通常比通用搜索提供更高质量的结果,但为给定查询选择最优智能体仍是未解决的问题。现有方法基于推断的主题或意图路由查询,但基于意图的选择存在根本局限:它未纳入检索内容的信号,且无法检测到主题匹配的智能体产生低相关性结果的情况。我们通过训练小型语言模型(SLM)解决该问题,先进行监督微调(SFT)再开展强化学习,使其能联合执行智能体选择与下游工具调用的结构化参数生成,采用基于检索相关性及查询-智能体主题对齐的分层奖励函数。这使模型能从检索性能中学习任务相关的智能体适用性:哪些智能体对哪些查询分布能可靠产生高相关性结果,以及何时尽管存在表面主题重叠仍将查询从专用智能体重定向。在这类智能体-查询不匹配的目标子集上,训练后的模型NDCG@10达0.918,而仅基于意图路由的两个大语言模型(LLM)基线Amazon Nova Lite和Claude Haiku 4.5分别为0.539和0.490;总体而言,其平均NDCG@10达0.771(较Nova Lite提升0.177,较Haiku提升0.219),平均选择延迟为120.1ms,较Nova Lite降低82.4%。
英文摘要
Specialised retrieval agents typically surface higher quality results than general-purpose search, but selecting the optimal agent for a given query remains an open problem. Current approaches route queries based on inferred topic or intent, however intent-based selection is fundamentally limited: it does not incorporate signal from retrieved content, and cannot detect when a topically aligned agent produces low-relevance results. We address this by training a small language model via supervised fine-tuning followed by reinforcement learning to jointly perform agent selection and structured parameter generation for downstream tool calls, using a hierarchical reward function grounded in retrieval relevance along with query-agent topic alignment. This enables the model to learn task-dependent agent suitability from retrieval performance: which agents reliably yield high-relevance results for which query distributions, and when to redirect queries away from specialised agents despite surface-level topical overlap. On a targeted subset of such agent-query mismatches, the trained model achieves an NDCG@10 of 0.918 compared to 0.539 and 0.490 for two LLM baselines (Amazon Nova Lite and Claude Haiku 4.5) that route on intent alone. Overall, it achieves a mean NDCG@10 of 0.771 (+0.177 over Nova Lite, +0.219 over Haiku) with a mean selection latency of 120.1ms, an 82.4% reduction over Nova Lite.