arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17994cs.CL

判断、检索或弃权:带有可证明风险保证的不确定性引导大语言模型(LLM)判断

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

  • Dalhousie University(达尔豪斯大学)
  • New York University Abu Dhabi(纽约大学阿布扎比分校)
  • Emory University(埃默里大学)

机构由 AI 辅助整理,请以论文原文为准。

Sher Badshah, Ali Emami, Hassan Sajjad

AI总结:

本研究针对无参考LLM评判的可靠性问题,提出带可证明风险保证的不确定性引导双模式风险控制框架,在维持目标错误率的同时提升了判决覆盖率。

AI中文摘要:

将大语言模型(LLM)用作评判者已成为大规模评估模型输出的标准做法,这在评估有用性或对齐性等主观开放式任务中尤为常见,此类任务不存在单一参考答案。然而,对于无参考的LLM评判而言,客观任务会带来独特的可靠性挑战。在缺乏参考答案的情况下,评判者要么通过自身参数知识,要么通过工具增强来评估事实正确性。前者虽能实现高效评估,但评判者可能会产生幻觉或缺乏判决所需的足够证据;相反,工具增强可提供额外证据,但会引入额外计算成本,且需要合适的机制来确定何时以及如何可靠使用这些证据。更重要的是,两种方法均无法对已接受判决的风险提供形式化控制,也无法在指定水平上保证其可靠性。我们提出一种风险控制框架,该框架在保留集上校准不确定性阈值,使得在已接受判决中,错误发现率以高概率保持在用户指定水平α以下,采用有限样本Clopper-Pearson区间。当参数模式的置信度不足时,实例会被路由至检索增强模式,此时评判者会收集网络证据并在第二个校准阈值下重新评估该实例。该有限样本保证可推广到这种双阈值路由,无需额外假设。在开放域问答基准及不同规模的评判者上,该框架维持了目标错误率,同时实现了比单模式基线高得多的覆盖率。

英文摘要:

Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists. However, objective tasks introduce a distinct reliability challenge for reference-free LLM judging. In the absence of a reference answer, the judge evaluates factual correctness either through its parametric knowledge or through tool augmentation. Although the former enables efficient evaluation, the judge may hallucinate or lack sufficient evidence for its verdict. Conversely, tool augmentation can provide additional evidence but introduces extra computational cost and requires an appropriate mechanism to determine when and how that evidence should be used reliably. More importantly, neither approach alone provides formal control over the risk of accepted verdicts or guarantees their reliability at a specified level. We propose a risk-controlled framework that calibrates uncertainty thresholds on a held-out set so that the false discovery rate among accepted verdicts remains below a user-specified level~$α$ with high probability, using finite-sample Clopper--Pearson intervals. When the parametric mode is not sufficiently confident, the instance is routed to a retrieval-augmented mode, where the judge gathers web evidence and re-evaluates the instance under a second calibrated threshold. The finite-sample guarantee carries over to this two-threshold routing without additional assumptions. Across open-domain QA benchmarks and judges of varying scales, the framework maintains the target error rate while achieving substantially higher coverage than single-mode baselines.

补充信息

↑