发表机构
University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对金融文档问答中的令牌预算分配问题,提出按每令牌增益排序的选择性升级方法,在10%预算下提升裁决质量并节省46.6%令牌,发现成本可预测但收益排序仍是瓶颈。
AI 中文摘要
检索增强生成系统可以将困难查询路由到更深的上下文,但批量部署必须在各次调用之间分配共享的令牌预算,而各次调用的成本因查询而异。我们将选择性升级(selective escalation)形式化为面向金融文档问答的有限批次分配问题。150个FinanceBench问题中的每一个首先获得一个top-1检索答案。预测器估计可选top-5调用的裁决质量增益和令牌成本,分配器根据每个令牌的预测增益对调用进行优先级排序。在名义上的10%预算下,与仅按增益排序相比,按每令牌增益分配将裁决质量提高了0.034(95%文档自助法置信区间[0.001, 0.072]),同时使用的总令牌比一次性top-5检索少46.6%。额外调用的成本可准确预测(R平方为0.93),而有益的升级仍然难以排序(AUROC为0.60)。这些结果表明,在严格约束下,异构成本是可操作的,而整个预算前沿的进展取决于更强的查询特定收益估计。
英文摘要
Retrieval-augmented generation systems can route difficult queries to deeper context, but batch deployments must allocate a shared token budget across calls whose costs vary by query. We formulate selective escalation as finite-batch allocation for financial document question answering. Each of 150 FinanceBench questions first receives a top-1 retrieval answer. Predictors estimate the adjudication-quality gain and token cost of an optional top-5 call, and the allocator prioritizes calls by predicted gain per token. At the nominal 10% budget, gain-per-token allocation improves adjudication quality over gain-only ranking by 0.034 (95% document-bootstrap CI [0.001, 0.072]) while using 46.6% fewer total tokens than one-pass top-5 retrieval. Additional-call cost is accurately predictable (R-squared 0.93), whereas beneficial escalation remains difficult to rank (AUROC 0.60). These results show that heterogeneous cost is actionable under tight constraints, while progress across the full budget frontier depends on stronger query-specific benefit estimates.
Comments6 pages, 3 figures, 6 tables. Code and aggregate artifacts: https://github.com/junru-zhu/token-budgeted-escalation-financial-qa