发表机构
The University of Queensland; Polytechnic University of Milan(昆士兰大学; 米兰理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨多语言LLM路由中预生成成功探针的跨语言可靠性,发现英文训练探针迁移后校准度下降,而池化多语言监督提升判别力与校准度,使路由成功率提高0.7%并降低成本13.0%。
AI 中文摘要
预生成成功探针在解码前根据语言模型的隐藏激活估计响应正确性,从而实现成本感知的路由。虽然先前的工作主要证明了它们在英文输入上的实用性,我们沿三个维度研究了它们跨语言的可靠性:(1)它们是否保留可能成功与失败的排序(判别力);(2)它们是否保留与观察到的成功频率匹配的概率(校准度);(3)它们是否产生在候选模型间足够可比较的分数,以用于成本感知的多语言路由(实用性)。使用10种语言的3,000道MATH问题和8个开放权重模型配置,我们比较了从英文训练的探针和等预算的池化多语言探针的跨语言迁移。英文训练的探针保留了有用的跨语言判别力,但在迁移后校准度变差。池化的多语言监督改善了两个属性,并产生了更可靠的成功估计。在路由实验中,池化路由器实现了0.7%更高的测试成功率,同时相对于总是选择平均成功率最高的模型,将建模成本降低了13.0%。这些结果表明,多语言路由需要保持良好校准且跨语言和模型可比较的成功估计。
英文摘要
Pre-generation success probes estimate response correctness from a language model's hidden activations before decoding, enabling cost-aware routing. While prior work has demonstrated their utility primarily on English inputs, we study their reliability across languages along three dimensions: (1) whether they preserve the ranking of likely successes and failures (DISCRIMINATION); (2) whether they retain probabilities that match observed success frequencies (CALIBRATION); and (3) whether they produce scores comparable enough across candidate models for cost-aware multilingual routing (UTILITY). Using 3,000 MATH problems in 10 languages and 8 open-weight model configurations, we compare cross-lingual transfer from English-trained probes and equal-budget pooled multilingual probes. English-trained probes retain useful cross-lingual discrimination but become less well calibrated after transfer. Pooled multilingual supervision improves both properties and yields more reliable estimates of success. In routing experiments, the pooled router achieves a 0.7% higher test success rate while reducing modeled cost by 13.0% relative to always selecting the model with the highest average success. These results show that multilingual routing requires success estimates that remain well calibrated and comparable across languages and models.