思考成本即 token:更多结构何时值得付出代价
Thinking Costs Tokens: When More Structure is Worth the Price
- Royal Bank of Canada(加拿大皇家银行)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究语言模型推理结构的 token 预算阈值,在 FinQA、TAT-QA 任务上用 GPT-5.4 mini 验证,1500 token 后验证搜索架构性能优于单体模型,交叉点在 1000-1500 token 间。
AI中文摘要:
为语言模型添加推理结构可使其进行搜索、验证与修正,但这些操作会消耗其本应合理使用的预算。本文探究是否存在 token 预算阈值:低于该阈值时,规划与验证的开销会损害性能;高于该阈值时则会提升性能。我们在 FinQA 和 TAT-QA 金融推理任务上评估两个系统,使用 GPT-5.4 mini,覆盖 14 个预算层级,范围为 250 至 42000 个输出等效 token。第一个系统是单体模型,即单次 LLM 调用;第二个是验证搜索架构,添加了规划、无标签检查与修复能力。我们运行 1000 个案例,共完成 28000 个单元。两个系统在最低两个层级均得 0%,因无法容纳完整提示;在 1000 token 时,单体模型准确率达 18%,而验证搜索因规划开销无空间生成答案,得分接近 0%;从 1500 token 起,验证搜索超过单体模型并保持稳定优势,在最高层级约达 44%,单体模型约达 40%。交叉点位于 1000 至 1500 输出等效 token 之间,经严格的交并检验确认(两端 p ≤ 0.001)。
英文摘要:
Adding inference structure to a language model lets it search, verify, and revise, but these actions consume the very budget they are supposed to use well. In this paper, we investigate whether there exists a token-budget threshold, below which the overhead of planning and verification hurts performance and above which it helps. We evaluate two systems on FinQA and TAT-QA financial reasoning tasks, using GPT-5.4 mini across 14 budget tiers ranging from 250 to 42,000 output-equivalent tokens. The first system is a monolith, which is a single LLM call. The second is a verified search architecture that adds planning, label-blind checking, and repair capabilities. We run 1,000 cases for a total of 28,000 completed cells. Both systems score 0% at the two lowest tiers, where neither can fit a complete prompt. At 1,000 tokens, the monolith reaches 18% accuracy while verified search scores near 0%, since the planning overhead leaves no room for an answer. From 1,500 tokens onward, verified search surpasses the monolith and maintains a consistent advantage, reaching approximately 44% at the highest tiers while the monolith reaches approximately 40%. The crossover occurs between 1,000 and 1,500 output-equivalent tokens, confirmed by a strict intersection-union test ($p \le 0.001$ at both endpoints).