arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00106cs.LG

面向智能体工作流的组合式元路由学习:一个可执行基准

Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark

Natan Vidra, Alina Kapanova, Arun Kanhai, Spurthi Setty

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一个智能体工作流的可执行基准和感知预算的组合式元路由器,在保留测试集上实现100%成功率,成本比静态策略低43%,但在词汇偏移挑战任务上表现不佳,凸显词汇泛化是主要限制。

中文摘要 AI 辅助

智能体系统不仅需要决定输出什么答案,还需确定应执行哪些推理与执行操作,控制器可直接作答、分解请求、检索证据、执行代码、委派给专家或验证中间结果。现有路由研究大多单独选择模型端点、检索深度或工具。本文提出一个可执行基准和一个感知预算的元路由器,该元路由器可根据原始任务文本组合异构操作。该基准涵盖数据分析、冻结语料库研究和文档处理领域,包含216个训练任务、72个开发任务、108个保留测试任务及108个锁定词汇偏移挑战任务,操作执行后的结果由机器校验。本文采用独立正则化逻辑头,从词和字符特征中预测操作概率,在开发数据上进行温度缩放,并在路由成本和操作数量预算下贪婪地组合操作。在保留测试集上,所学策略成功率达100%,而强大的静态和固定工作流成功率为93.5%,成本比静态策略低43%;匹配的所学单步路由器成功率为56.5%。在未触及的挑战拆分上,所学成功率降至75.9%,落后于静态路由的93.5%,但成本仍比静态路由低49%,且比单步路由器高出34.3个百分点。该差距表明,词汇泛化而非路由执行是主要限制因素。这些结果建立了一个可复现的测试平台和一个有界的概念验证,并非是对活大型语言模型(LLM)性能的证明。

英文摘要

Agentic systems must decide not only what answer to produce, but which reasoning and execution operations should precede it. A controller may answer directly, decompose a request, retrieve evidence, execute code, delegate to a specialist, or verify an intermediate result. Existing routing work largely selects model endpoints, retrieval depth, or tools in isolation. We introduce an executable benchmark and a budget-aware meta-router that composes heterogeneous operations from raw task text. The benchmark contains 216 training, 72 development, 108 held-out test, and 108 locked lexical-shift challenge tasks across data analysis, frozen-corpus research, and document processing. Outcomes are machine checked after operations execute. Independent regularized logistic heads predict operation probabilities from word and character features, are temperature-scaled on development data, and are greedily composed under route-cost and action-count budgets. On the held-out test, the learned policy achieves 100% success versus 93.5% for strong static and fixed workflows, with 43% lower cost than the static policy; a matched learned one-shot router reaches 56.5%. On the untouched challenge split, learned success falls to 75.9% and trails static routing at 93.5%, while remaining 49% cheaper and exceeding one-shot routing by 34.3 points. The gap identifies lexical generalization, rather than route execution, as the principal limitation. These results establish a reproducible testbed and a bounded proof of concept, not evidence of live-LLM performance.

补充信息

↑