arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17719cs.SEcs.AIcs.CL

聚合分数所遗漏的:商业大语言模型API迁移中项目级回归的测量

What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

Xiaonan Xu, Wenjing Wu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对GPT-5.4到GPT-5.6 Sol的三次升级,发现聚合分数遗漏了项目级双向变化,发布了响应级档案与项目级评分输出。

中文摘要 AI 辅助

背景:依赖商业大语言模型(LLM)API的软件系统,在供应商弃用旧模型时必须迁移到后续版本。迁移决策通常依赖聚合基准分数,该分数将异构的项目级行为压缩为单一净数值。目的:本研究旨在测量这种压缩所掩盖的内容。方法:针对GPT-5.4到GPT-5.6 Sol产品序列中的三对升级,我们查询了900个公开基准项目(包括研究生水平知识、奥林匹克数学、指令遵循),每个项目在每个模型上查询50次,在错误发现率控制和实际显著性阈值下,将每个项目分类为可靠改进、可靠回归、实际等价或不确定,并针对标签排列空值对结果进行校准。结果:在所有9个迁移-基准组合中,可靠改进和可靠回归同时存在。聚合增益高达7.3个百分点的区间包含多达8.3%的可靠回归项目;聚合损失的区间包含多达10.7%的可靠改进项目。在指令遵循基准上,最新迁移中严格评分与宽松评分的差距扩大了3.9个百分点:严格评分下的3.9点回归在宽松评分下缩小至0.04点。结论:仅基于聚合分数的迁移决策会遗漏大量双向的项目级变化。完整的响应级档案和每个项目的评分输出已发布。

英文摘要

Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol product sequence, we query 900 public benchmark items (graduate-level knowledge, olympiad mathematics, instruction following) 50 times per item per model, classify each item as reliably improved, reliably regressed, practically equivalent, or inconclusive under false-discovery-rate control and a practical-significance threshold, and calibrate the results against a label-permutation null. Results: Across all nine migration-benchmark cells, reliable improvements and reliable regressions coexist. Edges with aggregate gains of up to 7.3 percentage points contain up to 8.3% reliably regressed items; edges with aggregate losses contain up to 10.7% reliably improved items. On the instruction-following benchmark, the gap between strict and loose scoring widens by 3.9 percentage points on the latest migration: a 3.9-point regression under strict scoring shrinks to 0.04 points under loose scoring. Conclusion: Migration decisions based on aggregate scores alone miss substantial bidirectional item-level change. The complete response-level archive and per-item scoring outputs are released.

发表机构

  • College of Computing, Georgia Institute of Technology(佐治亚理工学院计算学院)
  • University of Colorado Boulder(科罗拉多大学博尔德分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑