arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GenRubric:用于可扩展大语言模型评估的自进化评分规则生成

GenRubric: Self-Evolving Rubric Generation for Scalable LLM Evaluation

Yifan Chen, Haitao Li, Qingyao Ai, Fengbin Zhu, Tat-Seng Chua, Min Zhang, Yiqun Liu

arXiv 2608.29856首次发表:更新:

发表机构

Tsinghua University; National University of Singapore(清华大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GenRubric是一种自进化框架,通过强化学习结合多维度奖励优化评分规则生成,其在人工基准实验中提升了与专家评分的一致性,可泛化到新领域,助力可扩展的特定查询LLM评估。

AI 中文摘要

大型语言模型正日益被用作开放式任务的可扩展评估工具。然而,许多大语言模型评判在评分过程中会生成针对特定查询的标准,导致评估要求未被明确规定,且其覆盖范围难以审计。针对特定查询的评分规则可明确这些要求,但专家撰写的评分规则构建成本高昂,而现有自动方法通常依赖推理时的优化或外部监督。我们提出GenRubric,这是一种无需在自进化过程中额外人工标注即可从未标注查询中改进评分规则生成的自进化框架。我们的方法基于评分规则诱导的自一致性:针对同一查询独立采样的评分规则提供了其潜在评估要求的部分视角,而全面的评分规则应能生成在这些互补评估视角间具有泛化性的响应。我们通过强化学习实现这一原则,将跨评分规则的全面性信号与针对评分规则质量的组级和标准级奖励相结合。我们在多个领域训练了4B、8B和14B规模的GenRubric模型。在人工标注的评分规则基准上的实验表明,自进化提升了生成评分规则诱导的评估与专家撰写评分规则诱导的评估之间的一致性。该改进还可泛化到未见过的领域,证明了自进化评分规则生成在可扩展且针对特定查询的大语言模型评估方面的潜力。代码和模型可在该httpsURL公开获取。

英文摘要

Large language models are increasingly used as scalable evaluators for open-ended tasks. However, many LLM judges derive query-specific criteria during scoring, leaving the evaluation requirements insufficiently specified and their coverage difficult to audit. Query-specific rubrics make these requirements explicit, but expert-written rubrics are costly to construct, while existing automatic methods typically rely on inference-time refinement or external supervision. We introduce GenRubric, a self-evolving framework that improves rubric generation from unlabeled queries without requiring additional human annotations during self-evolution. Our approach is based on rubric-induced self-consistency: independently sampled rubrics for the same query provide partial views of its latent evaluation requirements, and a comprehensive rubric should induce a response that generalizes across these complementary evaluation views. We implement this principle through reinforcement learning, combining a cross-rubric comprehensiveness signal with group-level and criterion-level rewards for rubric quality. We train GenRubric models at 4B, 8B, and 14B scales across multiple domains. Experiments on human-annotated rubric benchmarks show that self-evolution improves the agreement between evaluations induced by generated rubrics and those induced by expert-written rubrics. The improvements further generalize to held-out domains, demonstrating the potential of self-evolving rubric generation for scalable and query-specific LLM evaluation. Code and models are publicly available at https://github.com/foggpoy/GenRubric.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑