arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过定价无标签检查实现风险可控的选择性大语言模型回答

Risk-Controlled Selective LLM Answering by Pricing Label-Free Checks

Dongyub Jude Lee, Jungseob Lee, Chanjun Park, Hyeonseok Moon, Heuiseok Lim

arXiv 2609.37493首次发表:更新:

发表机构

Zoom Communications; Korea University; Soongsil University(Zoom通讯公司; 高丽大学; 崇实大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PriceCheck通过为无标签检查定价并组合决策规则,在指定风险目标下选择性提供大语言模型答案,实现高覆盖率与低错误率,优于现有评分器。

AI 中文摘要

从大语言模型提供答案需要决定何时弃权(不执行),然而验证者的排序准确性本身并不能决定所提供答案中的错误率。我们引入了PriceCheck,它从无标签检查(如重新求解问题)中构建一个紧凑的决策规则族。每个检查都有一个价格:其在正确和错误答案上的一致率以及每次运行的成本。在小规模、类别增强的标注集上拟合的价格,组合成对调度覆盖率和成本的预测,指导运行哪些检查以及何时停止。随后,校准测试在指定的选择性风险目标下选择一个调度。在数学任务中,所选调度平均提供76.1%的答案,并在所有15个分割上保持留出集的选择性风险低于1.5%。在共享测试协议下,PriceCheck在该目标下提供的答案多于奖励模型、提示法官、生成器的置信度和训练过的正确性分类器。在匹配覆盖率下,它保持这些评分器中错误答案数量最少。在118个诊断调度中,基于价格覆盖率预测与观察覆盖率的秩相关系数为0.97。这些结果表明,选择如何组合和停止检查与验证者如何对答案排序同样重要。代码可在该https URL获取。

英文摘要

Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error rate among served answers. We introduce PriceCheck, which builds a compact family of decision rules from label-free checks such as re-solving a problem. Each check has a price: its agreement rates on correct and incorrect answers and its cost per run. Prices fitted on a small, class-enriched labelled set compose into predictions of a schedule's coverage and cost, guiding which checks to run and when to stop. A calibration test then selects a schedule at a stated selective-risk target. In mathematics, the selected schedules serve 76.1% of answers on average and keep held-out selective risk below 1.5% on all 15 splits. Under the shared testing protocol, PriceCheck serves more answers at that target than reward models, a prompted judge, the generator's confidence and a trained correctness classifier. At matched coverage, it keeps the fewest wrong answers among these scorers. Across 118 diagnostic schedules, price-based coverage predictions have a rank correlation of 0.97 with observed coverage. These results show that choosing how checks are combined and stopped matters alongside how well a verifier ranks answers. Code is available at https://github.com/js-lee-AI/PriceCheck.

Comments29 pages, 6 figures, 24 tables. Dongyub Jude Lee and Jungseob Lee contributed equally

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑