arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02632cs.LG

在线验证语言模型响应在成本约束下的方法

Online Verification of Language Model Responses Under Cost Constraints

发表机构加州大学尔湾分校
查看机构详情
  • University of California, Irvine(加州大学尔湾分校)

机构由 AI 辅助整理,请以论文原文为准。

Erfan Hajihashemi, Yanning Shen

首次发表
浏览论文内容

中文总结 AI 辅助

针对在线验证语言模型响应时成本与准确性的权衡,提出OMVV算法,通过维护多个弱验证器池并自适应路由,在有限时间内保证误接受与误拒绝率,同时实现次线性遗憾,实验证明其成本更低、准确性更高。

中文摘要 AI 辅助

随着大型语言模型越来越多地部署于多步推理任务中,验证其输出正确性对于维持大规模可靠性变得至关重要。验证大型语言模型输出的正确性通常通过查询代价高昂的ground-truth oracle来实现,但在在线设置中,每一步都调用该oracle是不切实际的。先前的工作通过在每个步骤查询单个弱验证器,并利用其分数来决定是否需要同时查询代价高昂的强验证器,从而将强验证仅保留给一小部分步骤。然而,单个固定的弱验证器可能无法随着输入查询的主题或难度随时间变化而保持一致的性能,并且提前承诺使用一个验证器可能导致验证成本过高或验证不准确。我们引入了OMVV(在线多验证器验证),一种算法,它维护一个包含K个候选弱验证器的池,这些验证器具有不同的成本和验证性能,并通过在线分数组合器和指数权重路由策略,自适应地将每轮决策路由到选定的验证器。OMVV提供了无分布假设的有限时间保证,涵盖整个验证器池的误接受率和误拒绝率,并在组合成本与一致性目标下,实现了相对于事后最优固定验证器的次线性遗憾。在推理数据集基准上的实验表明,在各种操作预算下,OMVV相比任何单一固定验证器,在更低验证成本下实现了更高准确性。

英文摘要

As large language models are increasingly deployed for multi-step reasoning, verifying the correctness of their outputs has become essential for maintaining reliability at scale. Verifying the correctness of large language model outputs is often done by querying a costly ground-truth oracle, which is impractical to invoke at every step in an online setting. Prior work addresses this by querying a single weak verifier on every step, and using its score to decide whether the costly strong verifier needs to be queried as well, reserving strong verification for only a small fraction of the steps. However, a single fixed weak verifier may not perform consistently well as the subject matter or difficulty of incoming queries changes over time, and committing to one in advance risks either overly costly or inaccurate verification. We introduce OMVV (Online Multi-Verifier Verification), an algorithm that maintains a pool of $K$ candidate weak verifiers with differing cost and verification performance, and adaptively routes each round's decision to a verifier selected via an online score combiner and an exponential-weights routing policy. OMVV provides a distribution-free, finite-time guarantee on false-accept and false-reject rates across the full pool of verifiers, and further achieves sublinear regret against the best fixed verifier in hindsight under a combined cost and consistency objective. Experiments on reasoning dataset benchmarks show that OMVV achieves higher accuracy at lower verification cost than any single fixed verifier, across a range of operating budgets.

↑