arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过动态评分标准共同进化大语言模型评估器和策略

Co-Evolving LLM Evaluators and Policies via DynamicRubric

Beining Wang, Weihang Su, Hongtao Tian, Hao Kong, Tao Yang, Ting Yao, Qingyi Pan, Yueyue Wu, Qingyao Ai, Min Zhang, Yiqun Liu

arXiv 2607.20083首次发表:更新:

发表机构

Tsinghua University; Tencent(清华大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于评估器反馈优化大语言模型时因策略改进导致的优化瓶颈问题,提出DynamicRubric框架,通过生成加权评分标准项实现评估器与策略共同进化,实验证明该框架提升了评估器性能及策略效果,并已成功应用于微信搜索。

AI 中文摘要

通过评估器对策略诱导样本的反馈进行训练后优化是改进大语言模型的主要机制。随着策略改进,采样响应质量趋同,导致策略优化瓶颈。本文从概率分配视角理论分析了差距的重要性,表明概率质量转移的方向增益即评估器分数差距,此为策略优化信号。基于此提出DynamicRubric框架,为候选集生成加权二元评分标准项并汇总为响应级分数。实验表明,DynamicRubric提升了评估器性能,优化的策略在可验证推理和编码任务中也有优势,还在微信搜索中部署并改善了关键指标。结果表明评估器应与所监督策略共同进化。

英文摘要

Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policies improve, these sampled responses become close in quality. These close candidates create a bottleneck for policy optimization: collapsed relative evaluator score gaps yield weak or misleading policy supervision. We theoretically characterize why these gaps matter through a probability allocation view, showing that the directional gain of shifting probability mass from one response to another is exactly the evaluator score gap between them. This identifies relative score gaps as the policy optimization signals that guide updates. Motivated by this view, we propose DynamicRubric, a response-set-conditioned evaluator--policy co-evolution framework that generates weighted binary rubric items for each candidate set and aggregates the resulting judgments into response-level scores. In our experiments with 8B backbones, DynamicRubric improves evaluator performance and provides stronger policy supervision than baselines using a 70B reward model or a 235B static rubric generator. DynamicRubric-optimized policies also show gains on verifiable reasoning and coding tasks. A DynamicRubric-optimized model is fully deployed in WeChat Search's AI answering scenario, where it serves all online traffic across tens of millions of requests per day and improves key online metrics. These results suggest a principle for evaluator-guided post-training: evaluators should evolve with the policies they supervise.

Commentsadd online model info (approved)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑