arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2608.14641cs.AIstat.ML

任务与会话级模型路由:四个基准测试中四种开源路由的通用接口混合评估

Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks

  • Indiana University(印第安纳大学)

机构由 AI 辅助整理,请以论文原文为准。

Kiran N. Kumar, Santhosh K. Saminathan

AI总结:

该研究提出通用测量协议,在四个基准上混合评估四种开源路由,发现多数路由层级分配恒定,vLLM语义路由随提示变化但无最高成功率,需控制固定层级基线以评估路由。

AI中文摘要:

智能体系统越来越多地将模型选择委托给路由组件,但开源路由通常在不同的任务、候选池和执行协议下进行评估,这限制了直接比较。我们提出了一种通用测量协议,并在RouterBench、BFCL v4、tau2-bench和WebArena四个基准上对四种路由实现进行混合评估。我们评估了290个冻结任务,对应锁定的2610个候选结果矩阵。三种路由会发出恒定或接近恒定的层级分配;只有vLLM语义路由会根据提示内容发生实质性变化,且它在四个基准中均未达到最高观测成功率。Always-Mid在三个基准上与Aurelio完全匹配,在第四个基准上误差为0.003以内。对于vLLM,任务级优势测试未检测到与匹配份额的内容盲分配相比的特定任务优势;仅在WebArena上以协议规定的5个百分点的边际建立了等价性。结果表明,在这些配置和控制下,观测到的增益与所选层级组成的关联比已证明的特定任务定位更紧密。因此,固定层级基线和所选层级分布是路由评估的必要控制项;该发现仅适用于这些配置、候选池和冻结基准样本,而非一般路由范式。

英文摘要:

Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually evaluated with different tasks, candidate pools, and execution protocols, limiting direct comparison. We present a common measurement protocol and hybrid evaluation of four router implementations across RouterBench, BFCL v4, tau2-bench, and WebArena. We evaluate 290 frozen tasks against a locked matrix of 2,610 candidate outcomes. Three routers emit constant or near-constant tier assignments; only vLLM Semantic Router varies materially with prompt content, and it has the highest observed success rate on none of the four benchmarks. Always-Mid matches Aurelio exactly on three benchmarks and within 0.003 on the fourth. For vLLM, task-level superiority tests detect no task-specific advantage over a share-matched content-blind allocation; equivalence is established only on WebArena at the protocol-declared five-percentage-point margin. The results show that, under these configurations and controls, observed gains track selected-tier composition more closely than demonstrated task-specific targeting. Fixed-tier baselines and selected-tier distributions are therefore necessary controls in router evaluation; the findings are scoped to these configurations, candidate pool, and frozen benchmark samples, not to routing paradigms in general.

补充信息

↑