arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07354cs.AI

评估用于LLM路由的升级信号:目标、控制以及五种自欺方式

Evaluating Escalation Signals for LLM Routing: Targets, Controls, and Five Ways to Fool Yourself

Ramin Pishehvar, Andrea Morandi, Mahesh Viswanathan

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一套检查清单,用于评估LLM路由中升级信号的可靠性,通过控制变量和难度基线等方法,防止将虚假信号误报为真实改进。

中文摘要 AI 辅助

决定何时将查询从小语言模型升级到大语言模型,需要一种廉价信号,该信号能在调用大模型之前预测升级是否有帮助。语义熵最初为检测幻觉而开发,是一个自然候选:它衡量模型采样答案在语义上的不一致程度,高不一致通常预示着不可靠的答案。我们在三个基准和两个模型家族上对其进行了测试。在GSM8K上,使用大小约相差十二倍的小/大模型对,语义熵能可靠地区分小模型的错误(AUROC 0.871),并在匹配成本下将路由准确率比随机升级提高多达九个百分点。一个在合成基准上早期看似强劲的结果被证明具有误导性:一个仅基于问题难度、不涉及模型的简单规则,几乎与语义熵完全匹配。本文的主要贡献是一组检查,可在结果被报告为真实之前发现此类问题。我们表明,将廉价的、仅基于问题的难度估计与任何信号一起评分,可以揭示该信号是增加了真实信息还是仅仅追踪问题看起来的难度;两种合理的“升级有效”定义在同一数据上可能产生截然不同的结果;基准可能几乎没有空间让任何信号胜过简单地始终使用大模型;并且实时采样的真实成本可能使路由比直接调用大模型更昂贵。对于重用缓存过去结果的更便宜替代方案,我们展示了如何预测其在新数据集上是否有效——通过提前正确预测从AUROC 0.908崩溃到随机水平(0.518)得到证实。我们提供这些作为评估升级信号的通用检查清单。

英文摘要

Deciding when to escalate a query from a small language model to a larger one requires a cheap signal that predicts, before the large model is called, whether escalating would help. Semantic entropy, originally developed to detect hallucinations, is a natural candidate: it measures how much a model's sampled answers disagree in meaning, and high disagreement often signals an unreliable answer. We test it across three benchmarks and two model families. On GSM8K, with a small/large pair about twelve times apart in size, semantic entropy reliably distinguishes the small model's mistakes (AUROC 0.871) and improves routed accuracy over random escalation by up to nine points at matched cost. An earlier strong-looking result on a synthetic benchmark proved misleading: a simple rule based only on question difficulty, with no model involved, matched semantic entropy almost exactly. This paper's main contribution is a set of checks that catch this before it is reported as real. We show that scoring a cheap, question-only difficulty estimate alongside any signal reveals whether the signal adds real information or just tracks how hard a question looks; that two reasonable definitions of "escalation worked" can produce very different results on the same data; that a benchmark can leave almost no room for any signal to beat simply always using the large model; and that the true cost of live sampling can make routing more expensive than calling the large model directly. For a cheaper alternative that reuses cached past outcomes, we show how to predict whether it will work on a new dataset -- confirmed by correctly forecasting a collapse from AUROC 0.908 to chance level (0.518) ahead of time. We offer these as a general checklist for evaluating escalation signals.

发表机构

  • Cisco Systems, Inc.(思科系统公司)

机构由 AI 辅助整理,请以论文原文为准。

↑