ARISE-RL:基于智能体评分准则的强化学习迭代自进化框架
ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对强化学习训练开放式智能体的奖励与评分准则难题,提出ARISE-RL框架,结合Generator与Solver的协同进化及RG-SED机制,在ECR-Bench上实现最优性能。
中文摘要 AI 辅助
通过强化学习(RL)训练开放式智能体面临两大挑战:一是缺乏可验证的标准答案及可扩展的评分准则;二是即便接近模型能力边界,长时序开放式智能体任务的奖励仍常表现出脆弱性与不稳定性,导致回合对比效果薄弱或存在噪声,掩盖了群体策略学习所需的细粒度优化信号。为解决这些问题,我们提出ARISE-RL,这是一种新型全周期自进化框架,通过评分准则介导的协同进化将任务/评分准则生成器(Generator)与推理求解器(Solver)耦合起来。Generator将与工具相关的评分准则建立在真实工具观测的基础上,其奖励来自生成与Solver不断演化的能力边界相匹配的有效中等难度任务;Solver则通过多步推理与工具使用,从细粒度的评分准则满足信号中学习。我们进一步提出奖励门控自进化蒸馏(RG-SED),仅当记忆增强的同策略变体产生经验奖励提升时,才将其选择性蒸馏回原策略,以此减少分布不匹配问题,避免对噪声指导的盲目模仿。最后,为支持严格评估,我们推出ECR-Bench,这是一套经专家校准的评分准则基准套件,涵盖单工具深度研究与多工具行程规划任务。大量实验表明,ARISE-RL在所有评估基准上均持续实现稳健且稳定的整体最优性能。
英文摘要
Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability boundary, long-horizon open-ended agentic tasks often yield brittle and unstable rewards, resulting in weak or noisy rollout contrast that obscures fine-grained optimization signals for group-based policy learning. To address these challenges, we propose ARISE-RL, a novel full-cycle self-evolution framework that couples a task/rubric Generator and a reasoning Solver through rubric-mediated co-evolution. The Generator grounds tool-related rubric criteria in real tool observations and is rewarded for producing valid, intermediate-difficulty tasks aligned with the Solver's evolving capability boundary. The Solver, in turn, learns from fine-grained rubric satisfaction signals through multi-step reasoning and tool use. We further introduce Reward-Gated Self-Evolution Distillation (RG-SED), which selectively distills a memory-augmented variant of the same policy back into itself only when the memory yields empirical reward improvement, thereby reducing distribution mismatch and avoiding blind imitation of noisy guidance. Finally, to support rigorous evaluation, we present ECR-Bench, an expert-calibrated rubric benchmark suite covering single-tool deep research and multi-tool travel planning. Extensive experiments demonstrate that ARISE-RL consistently achieves robust and stable overall state-of-the-art performance across all evaluated benchmarks.
发表机构
- Alibaba ATH Token Foundry(阿里巴巴ATH代币铸造厂)
- Hema, Alibaba Group(阿里巴巴集团盒马)
机构由 AI 辅助整理,请以论文原文为准。