发表机构
IBM Research India; IBM Research Almaden(IBM印度研究院; IBM阿尔马登研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出BenchDrift方法,通过生成基准问题的保义变体,发现大语言模型性能存在双向漂移,且漂移与重新措辞相关,模型性能提升未消除措辞敏感性。
AI 中文摘要
基准测试分数源于每个问题的单一措辞,这种单一措辞被当作能代表同一问题所有可能提问方式的整体空间,但实际并非如此。我们发现,在保持问题含义和答案固定的情况下重新措辞,会常规性地双向翻转模型的答案,导致一些失败变为成功,一些成功变为失败,我们将这种现象称为漂移。BenchDrift 沿着四个轴生成基准测试问题的保义变体,即语言轴、指称轴、语用轴和结构轴,并测量每种情况下正确性翻转的频率及原因。在八个模型和三个基准测试(GSM8K、MMLU、MATH-Hard)上,我们观察到漂移在双向都很显著。有两项关键发现:第一,措辞敏感性不会随模型性能提升而消失,反而会改变方向,弱模型从重新措辞中获得的收益多于损失,而强模型的损失远多于收益,因此基准测试中表现最佳的模型,其分数最依赖于所给的特定措辞;第二,尽管模型的漂移程度不同,但它们在很大程度上对哪些重新措辞会导致最多正确答案丢失达成一致,因此脆弱性属于重新措辞方式而非模型本身。此外,无论问题变短还是变长,重新措辞都会破坏模型原本有信心的答案。代码和数据:此 https URL
英文摘要
A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while keeping its meaning and answer fixed routinely flips a model's answer in both directions, so some failures become successes and some successes become failures. We call this drift. BenchDrift generates meaning-preserving variations of benchmark problems along four axes, namely linguistic, referential, pragmatic, and structural, and measures how often, and why, correctness flips under each. Across eight models and three benchmarks (GSM8K, MMLU, MATH-Hard), we observe that drift is large in both directions. Two findings stand out. First, phrasing sensitivity does not fade as models get better. Instead, it changes sign. Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain. We find that the best models on a benchmark are therefore the ones whose scores depend most on the wording they happened to be given. Second, the models largely agree on which rephrasings cost the most correct answers even though they differ in how much they drift, so fragility belongs to the rephrasing and not to the model. Furthermore, rephrasing breaks answers a model was confident about, whether the problem is made shorter or longer. Code and Data: https://github.com/IBM/BenchDrift/tree/demo-ui