LLM Router: 重新思考预填激活的路由
LLM Router: Rethinking Routing with Prefill Activations
浏览论文内容
中文总结 AI 辅助
本文提出基于预填激活的路由方法,通过分离编码器和目标模型,提升路由性能,实验显示其在成本和准确性上均优于传统方法。
中文摘要 AI 辅助
LLMs通常在平均基准准确度上表现相似,但在不同查询子集上表现出互补优势,表明具有查询特定模型选择的路由可超越单一模型。现有路由依赖语义查询特征,但难以捕捉模型特定失败或任务内在难度。我们研究通过内部预填激活进行路由。我们的关键思想,编码器-目标解耦,将生成预测信号的模型(编码器)与估计正确性的模型(目标)分离,使开放权重编码器可预测封闭源目标模型的性能。我们评估逐层几何探针,发现Fisher分离性(J)有效识别信息层,受有效维度性(d_eff)诊断支持。随后利用SharedTrunkNet,一种联合多输出MLP,通过连接的预填特征预测候选模型的同时正确性概率。在实验中,SharedTrunkNet一致优于语义基线。最佳情况下,SharedTrunkNet缩小了最强独立模型与oracle之间的45.58%差距,同时相对于最贵模型实现74.31%的成本节省。这些结果表明预填激活提供了一种稳健的路由信号,证明了机理路由作为高绩效替代纯语义选择的可行性。
英文摘要
Existing routers rely on semantic query features or handcrafted features, which often fail to capture model-specific failures or intrinsic task difficulty. We instead route using internal LLM activations, specifically the residual stream. Our key idea, Encoder-Target Decoupling, separates the model that produces the predictive signal (the Encoder) from the model whose correctness is being estimated (the Target), allowing open-weight encoders to predict the performance of closed-source target models. We evaluate layerwise geometric probes, finding that Fisher Separability ($J$) effectively identifies informative layers, supported by Effective Dimensionality ($d_{\mathrm{eff}}$) diagnostics. We then utilize a SharedTrunkNet, a joint multi-output MLP that predicts simultaneous correctness probabilities across candidate models using concatenated prefill features. In our experiments, SharedTrunkNet consistently outperforms semantic baselines. At its best, SharedTrunkNet closes 45.58% of the gap between the strongest standalone model and the oracle while achieving 74.31% cost savings relative to the most expensive model. These results demonstrate that prefill activations provide a robust routing signal, establishing activation-based routing as a high-performance alternative to purely semantic selection.
发表机构
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。