发表机构
Tsinghua University; Tencent(清华大学; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出EnGRICH框架,利用人类批评训练元批评器为生成式奖励模型提供过程监督,提升批评可靠性,在七个基准上超越现有方法。
AI 中文摘要
生成式奖励模型(GRM)对大型语言模型(LLM)的优化至关重要。与标量奖励模型不同,GRM在提供偏好判断的同时生成自然语言批评,从而提供更细粒度的评估信号。其有效性在很大程度上取决于批评的可靠性。然而,现有的GRM训练通常将最终偏好正确性作为结果监督。由于偏好结果空间高度受限,不可靠的批评仍可能产生正确结果,从而被强化。近期工作利用人类批评进行过程监督,但此类批评稀缺且常被简化为标量奖励,导致其细粒度评估信息未被充分利用。我们认为,从人类批评中学到的评估标准可以推广到更广泛的结果偏好数据。为此,我们提出EnGRICH,一种GRM训练框架,将GRM与从少量人类批评中训练的训练时元批评器(MetaCritic)配对。MetaCritic构建响应特定的评分标准,并利用这些标准评估生成批评的证据覆盖度和正确性。由此产生的信号既提供了细粒度信用分配的过程奖励,也为探索更优批评提供了结构化指导。在GRM训练过程中,MetaCritic进一步优化,以将人类基础的评估标准推广到结果数据。在推理时,训练好的GRM独立运行。在七个奖励模型基准上的实验表明,EnGRICH持续优于竞争基线,进一步分析验证了其核心机制的有效性。
英文摘要
Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside preference judgments, providing finer-grained evaluation signals. Their effectiveness depends heavily on critique reliability. However, existing GRM training typically uses final preference correctness as outcome supervision. Because the preference outcome space is highly constrained, unreliable critiques can still yield correct outcomes and thus be reinforced. Recent work leverages human critiques for process supervision, but such critiques are scarce and are often reduced to scalar rewards, leaving their fine-grained evaluative information underutilized. We argue that evaluative criteria learned from human critiques can be generalized to broader outcome-only preference data. To this end, we propose \textbf{EnGRICH}, a GRM training framework that pairs the GRM with a training-time MetaCritic learned from a small set of human critiques. MetaCritic constructs response-specific rubrics and uses them to evaluate the evidence coverage and correctness of generated critiques. The resulting signals provide both process rewards for fine-grained credit assignment and structured guidance for exploring better critiques. During GRM training, MetaCritic is further optimized to generalize human-grounded evaluative criteria to outcome-only data. At inference, the trained GRM operates independently. Experiments across seven reward-model benchmarks show that EnGRICH consistently improves over competitive baselines, while further analyses validate the effectiveness of its core mechanisms.