arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SchemaRouter:面向高效异构智能体检索增强生成的字段感知工具路由

SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAG

Yong-eun Cho

arXiv 2608.21375首次发表:更新:

发表机构

KailosLab(凯洛斯实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SchemaRouter是轻量级字段感知工具路由层,通过模式图实现工具与字段的精准选择,在材料科学基准中,其效率、准确率、工具精确率等表现优于基线,还能提供来源与许可信息。

AI 中文摘要

异构智能体检索增强生成(RAG)系统日益需要协调外部API、内部数据库、向量存储和图存储。将所有工具描述暴露给LLM智能体,或仅通过向量相似度选择工具,会导致两种代价高昂的故障:过度获取,会增加有效载荷大小、令牌使用量和延迟;以及获取不足,会遗漏回答查询所需的字段。我们提出SchemaRouter,一个轻量级路由层,它将工具、端点、参数、响应字段、领域概念、单位、来源和许可策略表示为模式图。给定一个查询,SchemaRouter会生成一个可执行的工具计划,指定要调用哪些工具以及要检索哪些字段。小型LLM会提取意图、概念和源约束,而字段选择则通过意图组投影和带有别名层的概念-字段匹配在图上确定性完成。在包含110个查询的材料科学基准测试中,SchemaRouter的答案准确率为0.71,与“全部获取”在重叠置信区间内匹配,且超过“全部提示”的0.66,尽管它们的区间存在重叠。它使用227个检索上下文令牌,而“全部获取”为2066个,并且比“全部提示”实现了2.7倍的端到端延迟降低。它还获得了0.93的最佳工具精确率和1.0的参数有效性。SchemaRouter在62%的答案中确立了来源和许可信息,而所有基线的这一比例约为0%。我们还发现,最小化所选字段数量会产生反效果:它将答案准确率降至0.56,令牌节省可忽略不计,而保留召回率的投影则恢复了最高准确率。SchemaRouter在保持竞争力准确率的同时,提高了效率、与模式大小无关的扩展性以及可验证的来源/许可基于的回答能力。

英文摘要

Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and graph stores. Exposing all tool descriptions to an LLM agent, or selecting tools only by vector similarity, causes two costly failures: over-fetching, which increases payload size, token use, and latency, and under-fetching, which omits fields needed to answer the query. We present SchemaRouter, a lightweight routing layer that represents tools, endpoints, parameters, response fields, domain concepts, units, provenance, and license policies as a schema graph. Given a query, SchemaRouter emits an executable tool plan specifying which tools to call and which fields to retrieve. A small LLM extracts intent, concepts, and source constraints, while field selection is deterministic over the graph through intent-group projection and concept-field matching with an alias layer. On a materials-science benchmark of 110 queries, SchemaRouter achieves answer accuracy of 0.71, matching fetch-everything within overlapping confidence intervals and exceeding prompt-all's 0.66, though their intervals overlap. It uses 227 retrieved-context tokens versus 2,066 for fetch-everything and achieves 2.7x lower end-to-end latency than prompt-all. It also obtains the best tool-exact rate of 0.93 and parameter validity of 1.0. SchemaRouter grounds provenance and license information in 62 percent of answers, compared with approximately 0 percent for all baselines. We also find that minimizing selected-field count is counterproductive: it reduces answer accuracy to 0.56 with negligible token savings, while recall-preserving projection restores top accuracy. SchemaRouter improves efficiency, schema-size-independent scaling, and verifiable provenance/license-grounded answering at competitive accuracy.

Comments13 pages, 4 figures, 6 tables. Code, benchmark, and fixtures: https://github.com/JDeun/SchemaRouter_research

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑