发表机构
University of Ottawa; Vector Institute(渥太华大学; 向量研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究如何用推理时信号预测LLM版本更新导致的样本级退化,对比单模型与跨版本信号,发现信号有效性具任务依赖性且无通用最优信号,部分跨版本信号可支持选择性回退,从业者可据此选择信号。
AI 中文摘要
前沿大语言模型(LLM)更新频繁,总体上通常优于前代模型,但总体收益对单个样本的参考价值有限:一次更新仍可能导致样本级退化,即旧模型下正确的响应在新模型下变为错误。本文研究如何利用推理时可用的信号预测此类退化。我们在统一增值测试中比较了单模型信号(置信度、对数几率间距、注意力熵)与跨版本信号(输出KL散度、似然漂移、 token级KL、表示漂移),该测试可分离每个信号相对于置信度基线的增益。在三个任务族(多项选择题回答MCQ、数学推理、代码生成)的六个基准数据集及六组模型更新对上,我们发现:(1)信号有效性具有任务依赖性:置信度在MCQ和较简单的数学任务上效果最强,而似然/KL信号在较难的数学和代码任务上增益最频繁;(2)没有任何信号在所有模型更新中是通用最优的;(3)即使置信度失效,部分跨版本信号仍保持信息性,且无需标签,这支持了选择性回退的概念验证,即把高风险样本路由回旧模型。从业者可利用这些任务级模式,针对特定更新选择应信任的退化预测信号。代码可在该https URL获取。
英文摘要
Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate. But aggregate gains say little about individual samples: an update can still cause sample-level regression, where a response correct under the old model becomes incorrect under the new one. This paper studies how to predict such regressions from signals available at inference time. We compare single-model signals (confidence, logit margin, attention entropy) against cross-version signals (output KL divergence, likelihood drift, token-level KL, representation drift) under a unified added-value test that isolates each signal's gain over a confidence baseline. Across six benchmarks in three task families (multiple-choice question answering, or MCQ; math reasoning; code generation) and six model update pairs, we find that (1) signal effectiveness is task-dependent: confidence is strongest on MCQ and simpler math, while likelihood/KL signals give the most frequent gains on harder math and code; (2) no signal is universally best across model updates either; and (3) some cross-version signals stay informative even when confidence fails, including without labels, which supports a proof-of-concept selective fallback that routes high-risk samples back to the old model. Practitioners can use these task-level patterns to choose which regression signal to trust for a given update. Code is available at https://github.com/jiashengsally/llm-regression-signals.