RecurSE:面向LLM规则评判器的有界递归自评估
RecurSE: Bounded Recursive Self-Evaluation for LLM Rubric Judges
浏览论文内容
中文总结 AI 辅助
本研究提出RecurSE,一种无需外部监督的LLM规则评判器优化方法,通过同步评判器与检查器的协同演化、接口解耦及PAV监控,在多基准上实现泛化提升,可增强下游策略对齐。
中文摘要 AI 辅助
LLM作为评判器对于评估开放式文本、指导后训练至关重要,但改进评判器通常依赖昂贵的标注、奖励模型或从更强教师模型的蒸馏。本研究在RL训练奖励中消除了外部黄金监督:模型自身的评估能力为其优化生成学习信号,即一种称为递归自评估(Recursive Self-Evaluation, RecurSE)的有界递归自我改进(RSI)闭环设置。我们研究两个核心问题:自我改进何时发生,以及何时必须停止?首先,RecurSE将可训练评判器在每条规则的 rubric(评分标准)下评估候选响应(第1轮),与同步的策略副本检查器配对,该检查器对照元评分标准审核评判器的推理,以提供标量过程奖励(第2轮)。为实现学习,接口解耦在结构上将检查器的标量分数与评判器的判定 token 隔离,消除了会夸大自分配奖励的退化 token 复制捷径。其次,由于无锚定的递归学习本质上是有界的,成对优势有效性(Pairwise Advantage Validity, PAV)作为无偏验证监控器,共同跟踪评判器准确率和检查器保真度,以可靠识别最优早停窗口。在Qwen3.5-9B、Gemma-4-E4B-it和Qwen3.6-27B上,RecurSE在保留的医学、成对、摘要和专业基准上实现了一致的泛化提升。 ablation( ablation 实验,即消融实验)表明,同步评判器-检查器协同演化优于冻结检查器、外部元评判器、自一致性和缩放教师蒸馏。此外,由我们的评判器策划的偏好对可有效增强下游策略对齐。因此,当显式解耦和监控自生成奖励的有效性时,面向LLM作为评判器的有界RSI是可行的。
英文摘要
LLM-as-judge is essential for evaluating open-ended text and steering post-training, yet improving the judge itself typically relies on expensive annotations, reward models, or distillation from stronger teachers. In this work, we eliminate external gold supervision from the RL training reward: the model's own evaluative capability generates learning signals for its optimization -- a closed-loop setting of bounded recursive self-improvement (RSI) termed Recursive Self-Evaluation (RecurSE). We study two central questions: when can self-improvement occur, and when must it stop? First, RecurSE pairs a trainable judge evaluating candidate responses under per-rule rubrics (Pass 1) with a synchronized policy-copy checker that audits the judge's reasoning against meta-rubrics to supply a scalar process reward (Pass 2). To enable learning, interface decoupling structurally isolates the checker's scalar score from the judge's verdict tokens, eliminating a degenerative token-copying shortcut that inflates self-assigned rewards. Second, because unanchored recursive learning is inherently bounded, Pairwise Advantage Validity (PAV) serves as an unbiased validation monitor that jointly tracks judge accuracy and checker fidelity to reliably identify the optimal early-stopping window. Across Qwen3.5-9B, Gemma-4-E4B-it, and Qwen3.6-27B, RecurSE achieves consistent generalization gains across held-out medical, pairwise, summarization, and professional benchmarks. Ablations demonstrate that synchronized judge-checker co-evolution outperforms frozen checkers, external meta-judges, self-consistency, and scaled teacher distillation. Furthermore, preference pairs curated by our judge effectively enhance downstream policy alignment. Bounded RSI for LLM-as-judge is thus viable when self-produced reward validity is explicitly decoupled and monitored.
发表机构
- College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
- Meituan LongCat Team(美团龙猫团队)
机构由 AI 辅助整理,请以论文原文为准。