arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniRRM:跨语言与评估范式的统一推理奖励模型

UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation Paradigms

Peng Lai, Yichao Du, Junchao Wu, Weibo Gao, Linan Yue, Longyue Wang, Weihua Luo, Derek F. Wong, Guanhua Chen

arXiv 2609.05910首次发表:更新:

发表机构

Southern University of Science and Technology; Alibaba Group; Wuhan University; University of Macau; University of Science and Technology of China; Southeast University(南方科技大学; 阿里巴巴集团; 武汉大学; 澳门大学; 中国科学技术大学; 东南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对开放式任务中奖励模型可靠性差、评估标准静态且多语言支持不足的问题,本文提出统一推理奖励模型UniRRM及多语言数据集MixReward,通过分阶段推理链动态生成标准,实现跨语言和多种评估范式下的高效判断。

AI 中文摘要

强化学习(RL)在具有可验证奖励的任务上表现出色,但在开放式任务中,奖励模型的可靠性仍然是一个关键挑战。现有解决方案要么依赖成本高昂的专有LLM-as-a-Judge系统,要么依赖缺乏可解释性的不透明标量奖励模型。近期关于生成式奖励模型的研究提供了一种有前景的替代方案,但这些模型仍受限于静态评估标准、碎片化的评估范式以及有限的多语言支持。为应对这些挑战,我们引入了\ extbf{MixReward},一个覆盖六个领域和103种语言的大规模多语言数据集,包含成对数据和列表式数据,并提出了\ extbf{UniRRM},一个支持多种语言和评估范式的统一推理奖励模型。UniRRM使用分阶段的推理链动态生成任务通用和指令特定的标准,从而实现细粒度、输入自适应的判断,同时保持跨语言的一致性。实验表明,UniRRM-8B和UniRRM-14B在多个基准上达到了与同规模最先进模型相当的性能,并且对未见过的评估范式有效。此外,消融研究验证了UniRRM的可靠性和有效性。

英文摘要

Reinforcement learning (RL) excels on tasks with verifiable rewards, but in open-ended tasks, the reliability of reward models remains a key challenge. Existing solutions either depend on costly proprietary LLM-as-a-Judge systems or opaque scalar reward models that lack interpretability. Recent works on generative reward models offer a promising alternative, but they remain constrained by static evaluation criteria, fragmented evaluation paradigms, and limited multilingual support. To address these challenges, we introduce \textbf{MixReward}, a large-scale multilingual dataset spanning six domains and 103 languages, containing both pairwise and listwise data, and propose \textbf{UniRRM}, a unified reasoning reward model supporting multiple languages and evaluation paradigms. UniRRM uses a staged reasoning chain to dynamically generate task-generic and instruction-specific criteria, enabling fine-grained, input-adaptive judgments while maintaining consistency across languages. Experiments demonstrate that UniRRM-8B and UniRRM-14B achieve performance close to the state-of-the-art for models of comparable size across multiple benchmarks, and are effective for unseen evaluation paradigms. In addition, ablation studies validate the reliability and effectiveness of UniRRM.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑