arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DRACO:面向长视界智能体训练的基于动态规则的细粒度信用分配

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

Shubham Gandhi, Saurabh Goyal, Kiran Kate, Yara Rizk

arXiv 2609.04094首次发表:更新:

发表机构

Carnegie Mellon University; IBM Research(卡内基梅隆大学; IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长视界智能体训练中缺乏可验证奖励的问题,提出DRACO方法,通过动态生成规则并重新分配判断结果,在AppWorld和Tau-Bench上均取得显著性能提升。

AI 中文摘要

可验证奖励强化学习在任务具备程序化检查器时效果良好,但大多数长视界智能体领域不存在此类检查器。我们在结果不可见的设定下开展研究,该设定中不存在真实成功信号。多准则规则是提供此类奖励的常用方式,每条轨迹仅评分一次,但单个标量信号在数十步的范围内表现不佳。我们提出DRACO:基于规则的优势分配用于信用优化,它在训练过程中动态生成规则以跟踪策略的不断变化的能力,每条完整轨迹对这些规则评分一次,并将该判断结果重新分配给负责标注规则的步骤,从而在GRPO中产生差异化的每步优势。该重新分配是闭式的,不引入任何训练过的归因模块。在AppWorld上,DRACO尽管自身未使用任何检查器,但相比基础模型获得了15.9分的提升,相比使用稀疏真实奖励训练的GRPO获得了5.3分的提升。在域外Tau-Bench上,即便没有前沿判断器,它相比基础模型也获得了5.3分的提升,优于真实奖励训练和其他基于规则的训练设定。DRACO的代码可在此URL获取。

英文摘要

Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popular way to supply such a reward; they are scored once per trajectory, but a single scalar is a poor signal across tens of steps. We propose DRACO: Distributing Rubric-based Advantage for Credit Optimization. It generates rubrics dynamically during training to track the policy's evolving capability, scores those rubrics once per completed trajectory, and redistributes that judgment over the steps responsible for annotated rubrics to produce differentiated per-step advantages in GRPO. The redistribution is closed-form and does not introduce any trained attribution module. On AppWorld, DRACO gains 15.9 points over the base model and 5.3 points over GRPO trained with a sparse ground-truth reward, despite not using any verifiers itself. On out-of-domain Tau-Bench, it gains 5.3 points over the base model even without a frontier judge, beating both ground-truth-reward training and other rubric-based training settings. The code for DRACO is available at https://github.com/IBM/draco.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑