发表机构
ShanghaiTech University; BIGAI(上海科技大学; 北京通用人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统探讨LLM-as-a-Judge在开放式任务中的评判质量与下游效用,发现两者并非总一致,协议设计影响显著,且评判者引导能有效将测试时计算转化为性能提升。
AI 中文摘要
LLM-as-a-Judge(大语言模型作为评判者)越来越多地被用于评估缺乏标准答案的开放式任务中的策略响应。现有工作通常直接将评判结果转化为策略训练中的奖励信号,对评判的内在质量关注有限,并且大多将评判者的使用限制在训练时的监督中。我们通过考察评判的生成方式和使用方式,系统性地研究了评判质量与下游效用。对于评判的生成,我们沿三个维度改变评判者协议:裁决粒度、批评使用和评估批处理。对于评判的使用,除了策略训练之外,我们将评判者扩展到测试时推理,包括Best-of-N选择、评判者引导的修订和束搜索。我们发现:(i)令人惊讶的是,评判质量与下游效用并不总是一致的。(ii)评判者协议设计显著影响内在评判质量和下游效用。(iii)评判者引导有效地将测试时计算转化为性能提升,且收益因推理策略而异。我们的结果呼吁对开放式任务中的LLM评判者进行多方面的评估,涵盖内在评判质量和下游效用。
英文摘要
LLM-as-a-Judge is increasingly used to evaluate policy responses on open-ended tasks that lack ground-truth answers. Existing work often directly converts the resulting judgments into reward signals for policy training, paying limited attention to intrinsic judgment quality and largely restricting the use of Judges to training-time supervision. We systematically investigate judgment quality and downstream utility by examining both how judgments are elicited and how they are used. For judgment elicitation, we vary the Judge protocol along three dimensions: verdict granularity, critique usage, and evaluation batching. For judgment usage, beyond policy training, we extend Judge to test-time inference through Best-of-N selection, Judge-guided revision, and beam search. We find that, (i) Surprisingly, judgment quality and downstream utility do not always align. (ii) Judge protocol design substantially affects both intrinsic judgment quality and downstream utility. (iii) Judge guidance effectively converts test-time compute into performance gains, with benefits varying across inference strategies. Our results call for a multifaceted evaluation of LLM Judges on open-ended tasks, encompassing intrinsic judgment quality, and downstream utility.