AI 中文总结
本文提出基准天花板问题:随着AI模型在现有评估套件上达到天花板性能,区分信号集中在需要精英专家判断的最难项目上,导致评估信号逐渐枯竭。通过形式化模型、平台数据分析和治理讨论,揭示人类判断的结构性稀缺及其对AI能力度量的影响。
AI 中文摘要
基准是衡量、比较和治理AI能力的主要工具。本文认为,前沿AI基准的有效性取决于其构建中嵌入的人类判断的质量,而这种质量在结构上是稀缺的,标准规模化叙事掩盖了这一点。随着基础模型在现有评估套件上接近天花板性能,区分信号集中在最难的基准项目上,正是那些需要精英专家判断设计的项目。我们称此为基准天花板问题:当模型饱和了简单的大多数项目时,评估信号逐渐耗尽,而由少数高水平专家评估者编写的困难尾部仍然是真正区分的唯一来源。本文分三步论证这一论点。首先,我们提出了基准信号贬值的正式模型。基准分数是潜在模型质量的公共信号,但其精度内生地依赖于基准有效性。随着前沿能力提升以及污染或策略优化增加,固定基准作为测量工具会贬值。模型表明,有效信号集中在困难尾部项目上,这些项目的替换成本随前沿能力凸性上升,且私人基准生产者相对于社会最优水平对有效性投资不足。其次,利用覆盖一千多名有资质专业人员的微1平台数据,我们记录了与高判断、低可编码性评估劳动相关的稀缺溢价。第三,我们探讨了政治经济学和治理影响。
英文摘要
Benchmarks are the primary instruments through which AI capability is measured, compared, and governed. This paper argues that the validity of frontier AI benchmarks is a function of the quality of human judgment embedded in their construction, and that this quality is structurally scarce in ways that standard scaling narratives obscure. As foundation models approach ceiling performance on existing evaluation suites, discriminating signal concentrates in the hardest benchmark items, precisely those requiring elite expert judgment to design. We term this the benchmark ceiling problem: the progressive exhaustion of evaluation signal as models saturate the easy majority of items while the difficult tail, authored by a thin stratum of highly expert evaluators, remains the only source of genuine discrimination. The paper develops this argument in three steps. First, we present a formal model of benchmark signal depreciation. Benchmark scores are public signals of latent model quality, but their precision depends endogenously on benchmark validity. As frontier capability rises and as contamination or strategic optimization increases, fixed benchmarks depreciate as measurement instruments. The model shows that valid signal concentrates in hard-tail items, that the replacement cost of such items rises convexly with frontier capability, and that private benchmark producers underinvest in validity relative to the social optimum. Second, drawing on platform data from micro1 covering over one thousand credentialed professionals, we document the scarcity premium associated with high-judgment, low-codifiability evaluation labor. Third, we develop the political economy and governance implications.