发表机构
University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文构建AI推理市场质量调整价格指数,发现当前方法漏测87%的价格降幅,买方价格按任务计算已停止下降,且AI评估的排行榜稳定性论点无法为相关经济统计辩护。
AI 中文摘要
自2024年以来,AI推理的挂牌价格稳步下降,但该降幅的测算几乎完全取决于测算方法。本文利用公开数据构建了AI推理市场的质量调整价格指数。该面板汇集了3208个模型、86家提供商的21024条挂牌价格观测值,并通过从基准响应模式估计的潜在质量指数将其与4605个基准分数关联,因此这里的享乐主义传统的质量阶梯是通过评估而非产品特征构建的。按照统计机构应用于软件的匹配模型方法测算,推理价格每年下降0.10个对数点。而质量调整后的指数每年下降0.73个对数点,因此87%的降幅是当前方法无法观测到的,这对该市场的测算竞争、集中度和生产率产生了直接影响。此外,按完成的任务计算,买方的价格停止下降。推理模型的令牌消耗速度快于令牌价格的下降速度,因此卖方价格与买方价格出现分化。一项预先注册的有效性审计对质量测算进行约束,得出最明确的结果。排除受污染标记的基准后,模型排名保持0.998的一致性,但指数每年变动0.49个对数点,因此AI评估中标准的排行榜稳定性论点无法为基于基准构建的经济统计提供辩护。价格、质量和审计均可从公开来源零成本完全复现。
英文摘要
Posted prices for AI inference have fallen steadily since 2024, yet the measured speed of that fall depends almost entirely on the method of measurement. This paper constructs quality-adjusted price indices for the AI inference market from public data. The panel assembles 21,024 posted-price observations across 3,208 models and 86 providers and joins them to 4,605 benchmark scores through a latent quality index estimated from benchmark response patterns, so the quality ladder of the hedonic tradition is built here from evaluations in place of product characteristics. Measured by the matched-model methods that statistical agencies apply to software, inference prices fell at 0.10 log points a year. The quality-adjusted index fell at 0.73, so 87% of the decline is invisible to current methods, with direct consequences for measured competition, concentration and productivity in this market. Counted per completed task, moreover, the buyer's price stopped falling. Reasoning models raised token consumption faster than token prices fell, and the seller's and buyer's prices accordingly diverged. A pre-registered validity audit disciplines the quality measure and yields the sharpest result. Excluding contamination-flagged benchmarks leaves model rankings intact at 0.998 yet moves the index by 0.49 log points a year, so the leaderboard-stability arguments standard in AI evaluation offer no defence of economic statistics built on benchmarks. Prices, quality and the audit are fully reproducible from public sources at zero cost.
Comments38 pages, 3 figures. Data, code and provenance archived at https://doi.org/10.5281/zenodo.22177190; pre-registration at https://doi.org/10.17605/OSF.IO/5UQJ2