arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型路由差距主要源于任务类型

Most of the LLM Routing Gap Is Task Type

Janghoon Lee

arXiv 2608.23023首次发表:更新:

发表机构

Redrob(雷德罗布)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究发现LLM路由的性能差距主要源于任务类型,提前为每种任务类型分配固定模型的静态策略,成本低于最佳单一模型,可优化大部分问题,而学习型路由器的优化目标问题数量少于运行间波动。

AI 中文摘要

大语言模型(LLM)路由器会选择由哪个模型回答每个查询,其吸引力在于不同模型对不同问题的表现存在差异。无论哪款单一模型整体表现最佳,仍会答错部分问题,而模型库中的另一款模型能答对其中很多问题。每次都选对模型的表现是该任务的上限,路由器正是为接近这一上限而设计的。然而,近期研究显示路由器并未接近这一上限:在五个基准测试上的21种路由方法中,差异显著的设计之间性能仅相差零点几分,且所有方法都远低于该上限;学习得到的路由器通常无法击败始终调用最强模型的简单策略。本研究探究这些未被路由正确处理的问题的共性:设置14款模型回答294个问题,涵盖韩语、英语、印地语三种语言的7种任务类型;对整个矩阵运行两次,未做任何修改,但4116个模型-问题对中有5.37%的得分出现差异,这种运行间波动是正常现象,且小幅度的性能提升并不能表明路由(无论是本研究的还是其他研究的)发挥了作用。仅当模型在两次运行中都答对时,答案才计为正确,据此该矩阵中有29个问题可通过路由优化;本研究中所有正确答案计数均遵循此规则。任务类型是这些可优化问题的主要来源:提前为每种任务类型分配一个模型(选定后不再更新),可优化29个问题中的21个;按语言拆分每种任务类型可再优化2个,剩余6个问题未优化。这少数问题本应是学习型路由器的优化目标,但其数量少于上述运行间波动的问题占比(波动是模型-问题对的占比而非问题的占比)。本研究采用的静态表策略以每次运行3.33美元的成本回答294个问题中的262个,而最佳单一模型以每次7.69美元的成本回答245个问题;所有这些均在相同的294个问题上拟合和评分,未使用保留集。

英文摘要

An LLM router picks which model should answer each query. The appeal is that models fail on different questions. Whatever single model is best overall still gets some wrong, and another model in the pool gets many of those right. Getting that choice right every time is the ceiling, and a router is an attempt to approach it. However, recent work reports that routers do not get close. Across 21 routing methods on five benchmarks, sharply different designs land within a fraction of a point of each other, and all of them stay far below that ceiling. Learned routers often fail to beat simply always calling the strongest model. We ask what those missed questions have in common. We set fourteen models to answer all 294 questions, with 7 task types across 3 languages: Korean, English and Hindi. We ran the whole matrix twice, changing nothing, but 5.37% of the 4,116 model-question pairs came out scored differently anyway. Run-to-run movement like that is normal, and we argue that a small win does not show that routing did anything, ours or anyone else's. Counting an answer correct only when the model got it right in both runs, 29 questions on this matrix can be improved with routing. Every correct-answer count here is on that rule. Task type accounts for most of them: assigning each task type one model in advance, chosen once and never updated, improves 21 of the 29. Splitting each task type by language improves 2 more and leaves 6 of 294 unoptimized. That handful is what a learned router would have been built for, and it is smaller than the run-to-run movement above, which is a share of pairs rather than of questions. The static table we adopted answers 262 of 294 questions at \$3.33 per run, against the best single model's 245 at \$7.69. All of this is fitted and scored on the same 294 questions with no holdout.

Comments21 pages, 2 figures, 12 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑