AI 中文总结
该研究提出令牌经济得分(TES)指标,通过151组实证分析发现任务结构比难度更影响推理效率,推理投入存在收益递减,应按需选择性启用大语言模型推理。
AI 中文摘要
仅针对推理型大语言模型的准确率基准测试忽略了一个核心部署问题:扩展推理令牌何时能抵消其成本?我们提出令牌经济得分(TES),这是一种边际基准测试指标,用于衡量推理模型相较于非推理基线模型的准确率提升,并按生成令牌乘数进行归一化处理。针对带推理切换功能的模型家族,我们定义了配对TES变体;针对无直接非推理对应模型的前沿模型,我们定义了近似TES变体。随后,我们在涵盖数学、代码生成、科学推理、指令遵循、专业知识、知识回忆及研究级物理的7个基准上,开展了151组模型-基准评估运行的实证基准分析。该分析考察了三个面向部署的维度:何种任务结构能产生正边际推理效率、在模型家族中增加推理投入如何改变TES、部署场景如何改变经济可行性。结果表明,任务结构比名义难度更能预测推理效率:AIME 2025和LiveCodeBench等序列推理链任务显示出高TES,而MMLU-Pro等知识回忆任务尽管难度较高,但TES较低。我们还发现,在更高推理投入水平下存在系统性的收益递减,包括额外思考会降低准确率的情况。最后,推理成本份额(RCS)显示,推理支出通常由内部思考主导;部署成本乘数(DCM)显示,本地部署可改变原本成本高昂的推理工作负载的经济性。这些发现支持了一种基准驱动的模型选择规则:应根据任务类型、投入水平和部署场景选择性启用推理,而非将其视为普遍有益的模式。
英文摘要
Accuracy-only benchmarking of reasoning-capable large language models misses a central deployment question: when do extended thinking tokens earn their cost? We introduce the Token Economy Score (TES), a marginal benchmarking metric that measures the accuracy gain of a reasoning model over a non-reasoning baseline, normalized by the generated-token multiplier. We define paired and approximated TES variants for model families with reasoning toggles and frontier models without direct non-reasoning counterparts. We then conduct an empirical benchmarking analysis across 151 model-benchmark evaluation runs on seven benchmarks spanning mathematics, code generation, science reasoning, instruction following, expert knowledge, knowledge recall, and research-level physics. The analysis examines three deployment-facing dimensions: which task structures yield positive marginal reasoning efficiency, how increasing reasoning effort changes TES within model families, and how deployment context changes economic viability. Results show that task structure predicts reasoning efficiency better than nominal difficulty: sequential inferencechain tasks such as AIME 2025 and LiveCodeBench show high TES, while knowledge-recall tasks such as MMLU-Pro show low TES despite their difficulty. We also find systematic diminishing returns at higher reasoning effort levels, including cases where additional thinking reduces accuracy. Finally, Reasoning Cost Share (RCS) shows that inference spend is often dominated by internal thinking, while Deployment Cost Multiplier (DCM) shows how on-premises deployment can change the economics of otherwise costly reasoning workloads. These findings support a benchmarking-driven model-selection rule: enable reasoning selectively by task type, effort level, and deployment context rather than treating it as a universally beneficial mode.
Comments15 pages, 8 figures and tables, accepted at the 2026 TPCTC Conference and will be published at Performance Evaluation and Benchmarking (Springer)