发表机构
University of Southern California; University of Chicago; University of California, Berkeley; Massachusetts Institute of Technology; Stanford University; University of California, Davis; Pennsylvania State University; Harvard University; University of Oxford(南加利福尼亚大学; 芝加哥大学; 加利福尼亚大学伯克利分校; 麻省理工学院; 斯坦福大学; 加利福尼亚大学戴维斯分校; 宾夕法尼亚州立大学; 哈佛大学; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大型语言模型进展的标量总结问题,通过LiveCodeBench基准分离混淆,发现2024年9月后发布的模型在竞赛编程难任务上有超出预期的增益,发布了相关数据集与代码。
AI 中文摘要
大型语言模型的进展常通过单一标量指标总结,如时间跨度、潜在能力估计或聚合基准分数,这些总结能捕捉整体性能,但未检验进展在不同任务难度上的分布是否存在差异。我们发现,大部分向更难任务增益的表观转移,并未反映难度-响应曲线形态的变化:在METR时间跨度数据上,一个包含能力提升的单Rasch模型可复现该模式,因此这主要由天花板效应而非能力的质性变化解释,这与指标选择可能使所谓涌现能力看似是模型本身属性的情况类似。随后我们识别出一种在该控制下仍存在的较小难度任务效应。在智能体基准上分离该效应较为困难,因为新模型通常搭配新的智能体框架运行,故难任务的增益无法归因于模型或其框架。我们通过LiveCodeBench解决这一混淆,这是一个公共竞赛编程基准,无需智能体框架,且将不同时间的模型与外生难度排序配对。在考虑整体能力提升后,2024年9月后发布的模型在最难问题上的增益,超出其简单和中等性能的预测值,在最保守假设下约为+0.40 logits,使难问题解决率从约18%提升至25%。该效应由最强推理模型主导,且适用于仅需短推理而非长期自主的难任务。我们将此视为竞赛编程特有的结果,因为我们的清晰识别依赖于单一编码基准。我们发布LiveCodeBench难度面板(66个不同时间的模型×1055个问题)及分析代码。
英文摘要
Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. We find that most of the apparent shift in gains toward harder tasks does not reflect a change in the shape of the difficulty-response curve. On METR time-horizon data, a single Rasch model with rising ability reproduces this pattern, so it is largely explained by ceiling effects rather than a qualitative change in capability. This echoes how the choice of metric can make claimed emergent abilities look like a property of the models themselves. We then identify a smaller hard-task effect that survives this control. Isolating it is difficult on agentic benchmarks, because newer models are usually run with newer agentic harnesses, so a gain on hard tasks cannot be assigned to the model or its scaffold. We break the confound with LiveCodeBench, a public competitive programming benchmark that runs no agentic scaffold while pairing dated models with an exogenous difficulty ordering. After accounting for the rise in overall ability, models released after September 2024 still gain on the hardest problems beyond what their easy and medium performance predicts, by about +0.40 logits under our most conservative assumption, raising the hard-problem solve rate from roughly 18% to 25%. The effect is led by the strongest reasoning models and holds for hard tasks that need only short reasoning, not autonomy over long horizons. We present this as a result specific to competitive programming, since our clean identification rests on a single coding benchmark. We release the LiveCodeBench Difficulty Panel (66 dated models x 1,055 problems) and our analysis code.
Comments25 pages, 4 figures, 7 tables. Data and code: https://github.com/harvenstar/CurveShift