Dr.Credit:面向深度研究智能体的基于评分标准的流程信用分配
Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
浏览论文内容
中文总结 AI 辅助
Dr.Credit提出基于评分标准的流程信用分配方法,通过任务要求监督中间工具调用,结合GRPO优势,在深度研究智能体上超越开放基线并媲美专有模型。
中文摘要 AI 辅助
基于评分标准的任务越来越多地通过强化学习(RL)来解决,其中评分标准分数被用作训练奖励。然而,这些奖励通常只监督最终答案,而不区分中间决策的贡献。许多现有的信用分配方法依赖于真实答案来定义流程奖励,这限制了它们在缺乏规范解决方案的开放式任务中的适用性。为解决这一局限性,所提出的基于评分标准的信用(rubric-grounded credit)将任务要求作为最终答案评估和流程监督的共享参考。工具返回的信息根据其为满足每条评分标准相对于该评分标准已接受支持的历史所提供的额外支持进行评估。通过参考这些历史,信用能够区分新支持与轨迹中已存在的证据,同时识别对每条评分标准的部分支持。本文使用基于评分标准的信用在深度研究智能体的强化学习框架中监督中间工具调用。由此产生的流程优势与GRPO结果优势相结合,以指导研究决策,同时保留对最终报告质量的监督。在四个领域内和领域外基准上的评估表明,本文方法在每项主要指标和子指标上均优于所评估的开放深度研究基线。同时,使用8B参数骨干网络,训练后的智能体在平均性能上可与所评估的前沿专有模型相媲美。进一步的分析表明,在有限的研究轮次预算下,证据采集更高效且报告质量更高,这推动了将基于评分标准的流程监督扩展到更广泛的基于评分标准的任务。
英文摘要
Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on ground-truth answers to define process rewards, limiting their applicability to open-ended tasks without canonical solutions. To address this limitation, the proposed rubric-grounded credit uses task requirements as a shared reference for final answer evaluation and process supervision. The information returned by tools is assessed for the additional support it provides toward satisfying each rubric relative to that rubric's history of accepted support. By referencing these histories, credit distinguishes new support from evidence already present in the trajectory while recognizing partial support for each rubric. Dr.Credit uses rubric-grounded credit to supervise intermediate tool turns in an RL framework for deep research agents. The resulting process advantages are combined with GRPO outcome advantages to guide research decisions while retaining supervision of final-report quality. Evaluations on four in-domain and out-of-domain benchmarks show that Dr.Credit outperforms the evaluated open deep research baselines on every primary metric and submetric. Meanwhile, with an 8B-parameter backbone, the trained agent achieves average performance competitive with the evaluated frontier proprietary models. Further analyses suggest more efficient evidence acquisition and higher-quality reports under limited research-turn budgets, motivating the extension of rubric-grounded process supervision to a broader range of rubric-based tasks.
发表机构
- University of Chinese Academy of Sciences(中国科学院大学)
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- Alibaba Group(阿里巴巴集团)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。