思考的代价:推理工作量作为模型特定的API契约
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract
浏览论文内容
中文总结 AI 辅助
该研究以Sonnet 5模型为对象,通过配对对比实验发现显式高推理工作量API契约的调用成本更高,但未检测到准确率差异,相关结论仅适用于特定模型、任务和日期。
中文摘要 AI 辅助
API购买者购买的是一份标注日期的契约,而非仅模型名称;该契约包含请求和提供的模型、推理工作量项(或其省略)、输出约束、服务产品、提示词以及价格表。我们通过对Sonnet 5模型进行注册配对对比研究,将其显式高推理工作量版本与省略推理工作量的同一模型进行对比,使用30道2026年美国数学邀请赛(AIME)题目,每道题调用5次。每次付费尝试被分配一个固定的终端类别,推理过程会对题目进行重采样,同时保留其重复调用。显式高工作量契约下的平均交付成本比省略工作量契约高0.01031美元/次调用,区间为[+0.00204, +0.01974];对应的准确率对比为+0.0133,区间为[-0.0267, +0.0467],我们未检测到准确率差异,该区间无法排除最高达4.67个百分点的准确率提升。注册点估计显示,高工作量契约下每道正确答案的成本为0.08665美元,省略工作量契约下为0.07662美元。标注日期的契约普查、Models-API元数据以及预注册的原始响应探针进一步记录了特定于模型的省略语义,包括同一提供商内部的情况;当原始结构不确定时,相关主张仍属于文档级别的范畴。请求注册表、解析器、终端分类法、统计计划和分析流程在检查结果前已被固定,所得主张仅适用于所研究的模型、任务和收集日期。
英文摘要
API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort omitted, using 30 AIME 2026 items and five calls per item. Every paid attempt was assigned one frozen terminal category, and inference resampled items while retaining their repeated calls. Mean delivered cost was \$0.01031 per call higher under the explicit-high contract than under the omitted contract [+\$0.00204, +\$0.01974]. The corresponding accuracy contrast was +0.0133 [-0.0267, +0.0467]; we did not detect an accuracy difference, and the interval permits a gain of up to 4.67 percentage points that this design cannot rule out. Cost per correct answer was \$0.08665 under the high-effort contract and \$0.07662 under the omitted contract, as registered point estimates. A dated contract census, Models-API metadata, and preregistered raw-response probes further documented model-specific omission semantics, including within a provider; claims remained at documentation grade when raw structure was indeterminate. The request registry, parser, terminal taxonomy, statistical plan, and analysis pipeline were frozen before outcomes were examined; the resulting claims are bounded to the model, task, and collection date studied.
发表机构
- Brandeis University(布兰戴斯大学)
- Brandeis School of Business and Economics(布兰戴斯商学院与经济学学院)
机构由 AI 辅助整理,请以论文原文为准。