发表机构
School of Artificial Intelligence, Beijing Normal University; School of Information Science and Engineering, Chongqing Jiaotong University; School of Computer Science and Technology, Dalian University of Technology; School of Computer Science and Technology, Jilin University(北京师范大学人工智能学院; 重庆交通大学信息科学与工程学院; 大连理工大学计算机科学与技术学院; 吉林大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出签名救援路由(SRR),通过分别预测大型模型纠正或破坏小型模型的结果,按差异排序请求,在固定预算下实现贝叶斯最优路由,实验证明优于传统不确定性路由。
AI 中文摘要
大型语言模型(LLM)级联使用小型模型回答简单请求,并将选定的请求升级到更大的模型。大多数路由器优先处理小型模型看起来不确定或可能出错的示例。这种代理忽略了一个决定性事实:只有当大型模型纠正了小型模型时,升级才有用;而当大型模型将正确答案替换为错误答案时,升级是有害的。我们引入了签名救援路由(SRR),一种预算路由方法,分别预测这两个事件,并根据它们的差异对请求进行排序。我们表明,在固定升级预算下,这种签名条件增益是贝叶斯最优路由评分。SRR在部署时仅需要小型模型的输出统计信息,并添加了一个轻量级的双头路由器。我们在MMLU、HellaSwag和ARC-Challenge的TBD示例上使用Qwen3-4B和Qwen3-8B评估了SRR。在准确率-计算曲线上,SRR达到了TBD的面积,而学习的小型模型错误预测器为TBD,熵路由为TBD。这些结果表明,预测增量价值而非模型不确定性,是高效LLM级联的一个简单而有效的目标。
英文摘要
Large language model (LLM) cascades answer easy requests with a small model and escalate selected requests to a larger model. Most routers prioritize examples on which the small model appears uncertain or likely to be wrong. This proxy ignores a decisive fact: escalation is useful only when the large model corrects the small model, and it is harmful when the large model replaces a correct answer with an incorrect one. We introduce Signed Rescue Routing (SRR), a budgeted routing method that predicts these two events separately and ranks requests by their difference. We show that this signed conditional gain is the Bayes-optimal routing score under a fixed escalation budget. SRR requires only the small model's output statistics at deployment and adds a lightweight two-head router. We evaluate SRR with Qwen3-4B and Qwen3-8B on TBD examples from MMLU, HellaSwag, and ARC-Challenge. Across the accuracy-compute curve, SRR reaches an area of TBD, compared with TBD for a learned small-model error predictor and TBD for entropy routing. These results show that predicting incremental value, rather than model uncertainty, is a simple and effective objective for efficient LLM cascades.