arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MatrixReward:基于评分矩阵的开放式生成奖励机制

MatrixReward: Reward from Rubric Matrix for Open-Ended Generation

Zihan Shen, Qi Liu, Zixuan Yang, Yiqun Chen, Chenglong Zhao, Xiaozhao Wang, Lei He

arXiv 2610.00389首次发表:更新:

发表机构

Zhejiang University; Renmin University of China; Qwen Business Unit of Alibaba(浙江大学; 中国人民大学; 阿里巴巴Qwen事业部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MatrixReward通过构建逐标准胜率矩阵并利用列离散度与相关性生成数据依赖的权重,结合理想轮廓距离计算奖励,在开放式问答基准上以63.02分超越基线约2%。

AI 中文摘要

开放式查询生成缺乏标准答案,因此需要有效的奖励机制。逐点评分标准在相同提示下仅提供样本答案相对质量的有限信息;将多个评分标准的判断合并为单一分数也可能掩盖这些答案之间的差异。我们提出MatrixReward,该方法通过在每个评分标准下比较每对采样响应,构建一个基于滚动的逐标准胜率矩阵来生成奖励。矩阵每列的离散度反映了该标准区分当前滚动的强度,而列间相关性则揭示了标准之间的重复性;这些统计量共同产生依赖于数据的标准权重。我们将这些权重与标准的先验权重相结合。经过列归一化和加权后,观测到的逐标准最大值和最小值定义了正理想轮廓和负理想轮廓。每个滚动到这两个理想轮廓的距离决定了其相对接近度的质量奖励。使用Qwen3-8B在四个开放式问答基准上进行评估,MatrixReward取得了63.02的平均分数,比最强基线高出约2.0%。这些结果支持了基于相对比较的矩阵可以更合理地用于开放式生成强化学习奖励构建的观点。

英文摘要

Open-ended query generation lacks standard answers, thus necessitating an effective reward mechanism. Pointwise scoring rubrics provide limited information about the relative quality of sample answers under the same prompt; merging multiple rubric judgments into a single score may also mask the differences between these answers. We propose MatrixReward, which constructs rewards from a rollout-by-rubric win-rate matrix obtained by comparing every pair of sampled responses under each rubric. The spread of each matrix column captures how strongly that rubric distinguishes the current rollouts, while correlations between columns reveal rubric repetition; together, these statistics yield data-dependent rubric weights. We combine these weights with the prior weights of rubrics. After column normalization and weighting, the observed per-rubric maxima and minima define positive and negative ideal profiles. Each rollout's distances to these two ideals determine its relative-closeness quality reward. Evaluated using Qwen3-8B on four open-ended query-answering benchmarks, MatrixReward achieves an average score of 63.02, outperforming the strongest baseline by approximately 2.0%. These results support the idea that matrices derived from relative comparisons can be used to construct rewards more reasonably for open-ended generative reinforcement learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑