arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.25598cs.AIcs.LG

面向过程监督的不可验证智能体任务的混合奖励归一化

Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks

发表机构阿里集团阿西莫团队 · 加州大学洛杉矶分校
查看机构详情
  • Accio Team, Alibaba Group(阿里集团阿西莫团队)
  • University of California, Los Angeles (UCLA)(加州大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

Peiran Xu, Zhuohao Li, Xiaoying Xing, Guannan Zhang, Debiao Li, Kunyu Shi

首次发表 更新
浏览论文内容

中文总结 AI 辅助

针对不可验证智能体任务中过程奖励难标注且易偏离最终结果的问题,提出 PPR 强化学习方法与 ReNorm 校准策略,在多类基准上达到最先进性能。

中文摘要 AI 辅助

大语言模型(LLM)越来越依赖搜索引擎等外部工具来解决需要推理和外部知识检索的复杂智能体任务。近年来,带可验证奖励的强化学习(RLVR)通过结果奖励对最终答案给予奖励,已证明其在提升 LLM 能力方面的有效性。尽管结果奖励易于监督,但只能提供稀疏信号和延迟反馈,这限制了其在长轨迹任务上的有效性。过程奖励通过评估中间步骤来解决这一问题,可提供细粒度监督并鼓励有据可依的问题求解。然而,逐步骤标注极其困难,尤其是在没有“黄金”答案的不可验证过程中。此外,逐步骤判断需要在局部质量与对最终结果的贡献之间取得平衡,因为朝更高过程奖励优化并不总能与更好的最终结果保持一致。为应对上述挑战,我们提出了 Principle Process Reward(PPR),这是一种将有原则的步骤级评估与结果验证统一起来的强化学习方法。我们训练了一个基于原则的奖励模型,以提高过程评估的透明度和可靠性,并进一步提出 Reward Normalization(ReNorm)策略来校准结果奖励与过程奖励。实验结果表明,PPR 在广泛的基准测试中取得了最先进性能,展现出显著的鲁棒性和泛化能力。我们的代码和模型集合可通过该链接获取。

英文摘要

Large Language Models (LLMs) increasingly rely on external tools such as search engines to solve complex agentic tasks that require reasoning and external knowledge retrieval. Recently, reinforcement learning with verifiable rewards (RLVR) has demonstrated its effectiveness in advancing capabilities of LLMs by rewarding the final answers via outcome rewards. While straightforward to supervise, outcome rewards only provide sparse signals and delayed feedback, which limits their effectiveness on long trajectories. Process rewards address this by evaluating intermediate steps, providing fine-grained supervision and encouraging grounded problem solving. However, it is notoriously hard to annotate step-wise labels, especially in non-verifiable process without "golden" answers. Furthermore, step-wise judgment requires the balance between local quality with contribution to the final outcome, as optimizing towards higher process reward may not always align with better final outcomes. To address the above challenges, we introduce Principle Process Reward (PPR), an RL approach that unifies principled step-level assessment and outcome verification. We train a principle-based reward model to improve the transparency and reliability of process evaluation, and further introduce a Reward Normalization (ReNorm) strategy to calibrate outcome and process rewards. Experiment results show that PPR achieves state-of-the-art performance across a wide range of benchmarks, demonstrating its impressive robustness and generalization. Our code and model collection is available in this link.

↑