arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29168cs.AI

JudgePanel:通过自适应多奖励强化学习实现面板审议的紧凑评判模型

JudgePanel: A Compact Judge with Panel Deliberation via Adaptive Multi-Reward Reinforcement Learning

  • AWS Generative AI Innovation Center(AWS生成式AI创新中心)

机构由 AI 辅助整理,请以论文原文为准。

Yiyue Qian, Shinan Zhang, Huan Song, Hannah Marlowe

AI总结:

本研究提出JudgePanel框架,通过自适应多奖励强化学习为紧凑评判模型配备多智能体面板审议能力,其推理成本仅为单模型水平,在四个评估基准上性能优于700亿参数的评判专用模型,可利用数百样本快速适配新领域。

AI中文摘要:

大语言模型作为评判者(LLM-as-a-Judge)范式已成为一种可扩展的人工评估替代方案。然而,单模型评判者受限于其固有的模型偏差,而通过多样化审议缓解该问题的多智能体评估协议在推理时成本过高。为此,我们提出JudgePanel,它为紧凑的评判(Judge)模型配备了多智能体面板(Panel)审议能力。具体而言,我们首先在由多个强评估器组成的集合生成的面板审议轨迹上进行训练,以捕捉讨论、分歧和解决的结构化模式。为进一步在监督微调(SFT)之外提升评判质量,我们引入AdaReward,这是一种自适应多奖励强化学习算法,在强化学习训练过程中,当不同目标以不同速率饱和时,动态重新平衡奖励组件的权重。为实现实际部署,我们还设计了一个轻量级领域专用模块,可利用数百个标注样本快速适配新的评估领域。结果表明:(i)新颖性:JudgePanel是首个为单个紧凑评判模型配备多智能体面板审议能力且推理成本仅为单模型水平的框架;(ii)有效性与可靠性:采用140亿参数主干网络的JudgePanel在四个评估基准上的性能优于参数达700亿的评判专用模型,表现出较强的位置一致性,且能利用数百个样本快速适配新领域。

英文摘要:

The LLM-as-a-Judge paradigm has emerged as a scalable alternative to human evaluation. However, single-model judges are limited by their inherent model biases, while multi-agent evaluation protocols that mitigate this through diverse deliberation are prohibitively expensive at inference time. To this end, we propose \textbf{\modelname}, which equips a compact \underline{Judge} model with multi-agent \underline{Panel} deliberation capability. Specifically, we first train on panel deliberation traces from an ensemble of strong evaluators, capturing structured patterns of discussion, disagreement, and resolution. To further improve judgment quality beyond SFT, we introduce \textit{AdaReward}, an adaptive multi-reward RL algorithm that dynamically rebalances reward component weights as different objectives saturate at different rates during RL training. For practical deployment, we further design a lightweight domain specialization module for rapid adaptation to new evaluation domains with few hundred labeled samples. As a result, (i) \textit{Novel}: the first framework to equip a single compact judge with multi-agent panel deliberation capability at single-model inference cost; (ii) \textit{Effective \& Reliable}: JudgePanel with a 14B backbone outperforms judge-specialized models up to 70B across four evaluation benchmarks, demonstrates strong position consistency, and rapidly specializes to new domains with few hundred samples.

补充信息

↑