arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BudgetVerify:面向金融问答的预算分层验证

BudgetVerify: Budget-Tiered Verification for Financial QA

Janet Jenq, Hongda Shen

arXiv 2609.33052首次发表:更新:

AI 中文总结

BudgetVerify提出预算分层验证框架,按需分配验证层级,在六个模型上实现更优的准确率-成本帕累托前沿,提升金融问答效率。

AI 中文摘要

金融问答通常需要精确的数字提取、单位处理以及对表格和文本的算术运算,但统一应用昂贵的验证会浪费测试时的计算资源。我们提出了BudgetVerify,一个预算分层的生成器-验证器框架,它将每个生成的答案路由到三个验证层级之一:不验证、轻量级检查与修订、或更高成本的先求解后比较验证。路由器基于离线正确性和令牌成本结果进行训练,在测试时,利用验证前可获得的信息(包括问题、上下文统计、生成的答案以及相关的生成器元数据)来选择验证层级。所选层级要么直接返回生成的答案,要么调用相应的验证器。在六个商业和开源基础模型上,BudgetVerify通过仅在有用时选择性地分配更强的验证,始终比固定验证策略产生更高效的准确率-成本帕累托前沿。尽管绝对性能因模型而异,但这些效率提升以及由此产生的定性前沿形状在不同生成器模型之间是一致的。

英文摘要

Financial question answering often requires precise numerical extraction, unit handling, and arithmetic over tables and text, but applying expensive verification uniformly wastes test-time compute. We propose BudgetVerify, a budget-tiered generator-verifier framework that routes each generated answer to one of three verification tiers: no verification, lightweight check-and-revise, or higher-cost solve-first-then-compare verification. The router is trained from offline correctness and token-cost outcomes and, at test time, selects a verification tier using information available before verification, including the question, context statistics, the generated answer, and associated generator metadata. The selected tier either returns the generated answer directly or invokes the corresponding verifier. Across six commercial and open-weight base models, BudgetVerify consistently produces more efficient accuracy-cost Pareto frontiers than fixed verification policies by selectively allocating stronger verification only when it is useful. Although absolute performance varies across models, these efficiency gains and the resulting qualitative frontier shape are consistent across generator models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑