测试时扩展:基于预算约束的多属性验证
Test-Time Scaling via Budgeted Multi-Attribute Verification
- City University of Hong Kong(香港城市大学)
- Hohai University(河海大学)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出BMA-GAI算法,在全局预算下联合选择候选答案与验证属性,通过成本感知臂选择和自适应采样实现高效认证,理论证明其一阶最优性,实验显示其优于现有方法。
AI中文摘要:
在共享计算预算下验证大语言模型生成的答案,需要联合决定检查哪些候选答案以及评估哪些验证属性。我们将此问题形式化为全局预算下的多属性好臂识别问题:每个候选答案是一个臂,沿多个代价高昂的属性进行评估,目标是尽可能多地认证那些在所有属性上平均得分超过规定阈值的候选答案。我们提出了BMA-GAI算法,该算法将成本感知的臂选择与属性的自适应采样相结合。每次观察既用于指导自适应分配,也用于支持随时有效的认证,从而无需单独的确认阶段。我们为BMA-GAI建立了渐近覆盖保证,并推导出一个匹配的信息论逆界,刻画了问题的内在复杂度,从而证明BMA-GAI在远离临界预算水平时是一阶最优的。在合成基准测试和LLM答案验证任务上的实验表明,与竞争方法相比,BMA-GAI能更高效地分配验证预算,并认证更多高质量候选答案。
英文摘要:
Verifying LLM-generated answers under a shared computational budget requires jointly deciding which candidates to inspect and which verification attributes to evaluate. We formulate this problem as multi-attribute good-arm identification under a global budget: each candidate is an arm evaluated along several costly attributes, and the goal is to certify as many candidates as possible whose mean scores exceed the prescribed thresholds on all attributes. We propose \textsc{BMA-GAI}, an algorithm that combines cost-aware arm selection with adaptive sampling of attributes. Every observation serves both to guide adaptive allocation and to support anytime-valid certification, which removes the need for a separate confirmation stage. We establish an asymptotic coverage guarantee for \textsc{BMA-GAI} and derive a matching information-theoretic converse that characterizes the intrinsic complexity of the problem, thereby proving that \textsc{BMA-GAI} is first-order optimal away from critical budget levels. Experiments on synthetic benchmarks and an LLM answer-verification task show that \textsc{BMA-GAI} allocates the verification budget more efficiently and certifies more high-quality candidates than competing methods.