发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究审计ChatGPT、Claude和Gemini,发现API基准分数在准确性和一致性上系统性地高于聊天界面,揭示API评估不能可靠代表部署系统的行为。
AI 中文摘要
基准分数是模型发布中的核心货币:它们影响购买决策、塑造公众信任并影响政策。然而,基准分数的一个关键假设是,通过API测量的模型性能能忠实反映部署系统的行为。我们通过审计ChatGPT、Claude和Gemini的七个系统和九个基准(涵盖通用能力、社会偏见和谄媚行为)来挑战这一假设。我们发现API与界面在准确性和一致性上存在系统性差异。平均而言,API评估的准确率得分比相应界面评估高3.4个百分点,重测一致性得分高2.1个百分点。对于ChatGPT,API与界面访问之间的性能差异超过了仅API访问下GPT 5.3与GPT 5.4之间的差异。换言之,切换访问界面可能造成的性能下降与降级整整一代模型相当。我们进一步测试了暴露的API控制(如改变系统提示、采样参数和推理设置)是否能复现界面行为。这些控制在某些情况下会改变行为,但不能可靠地消除差距。我们的发现记录了一个情境有效性差距:通过API获得的测量结果不一定能推广到相应的部署界面,这使得将API评估用作部署系统的代理变得复杂。
英文摘要
Benchmark scores are a central currency in model releases: they inform purchasing decisions, shape public trust, and influence policy. Yet, a key assumption underlying benchmark scores is that the model performance measured through APIs faithfully reflects the behavior of deployed systems. We challenge this assumption by auditing ChatGPT, Claude, and Gemini across seven systems and nine benchmarks spanning general capability, social bias, and sycophancy. We find systematic API--interface differences in both accuracy and consistency. On average, API evaluations score 3.4 percentage points higher in accuracy and 2.1 percentage points higher in test--retest agreement than corresponding interface evaluations. For ChatGPT, the performance difference between API and interface access rivals the API-only difference between GPT 5.3 and GPT 5.4. Put differently, switching access surfaces can degrade performance as much as downgrading a full model generation. We further test whether exposed API controls can reproduce interface behavior by varying system prompts, sampling parameters, and reasoning settings. These controls shift behavior in some cases but do not reliably eliminate the gap. Our findings document a context-validity gap: measurements obtained through APIs do not necessarily generalize to corresponding deployed interfaces, complicating the use of API evaluations as proxies for deployed systems.