发表机构
University of Illinois Urbana-Champaign; National University of Singapore; Zhejiang University; University of Rochester(伊利诺伊大学厄巴纳-香槟分校; 新加坡国立大学; 浙江大学; 罗切斯特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出依赖感知奖励塑形(DARS),通过依赖图分配步骤级信用,在多个任务上提升强化学习训练大型语言模型的成功率。
AI 中文摘要
在用强化学习训练大型语言模型时,终端奖励几乎无法提供关于哪些步骤重要的指导。常见的步骤信用分配方法忽视了这样一个事实:基于未纠正错误的工作是浪费的,而独立工作仍然有效。在仅有最终成功/失败奖励的情况下,失败回合中的每一步的总未来奖励为零,即使它取得了进展。我们提出依赖感知奖励塑形(DARS),该方法将任务进展表示为通过先决条件关系链接的谓词,并在依赖图上分配步骤级信用。一个标注器标记每个步骤验证、使无效或修复的谓词。已验证的谓词根据与最近损坏的先决条件的图距离进行折扣,而独立谓词不受影响。修复根据任何剩余错误更新这些权重;无效的谓词需要重新验证以重新获得信用。一个固定的势函数将这些标注转换为带符号的每步奖励。一个通用的奖励和标注接口允许DARS与一系列推理和智能体训练方法(如GiGPO和ARPO/AEPO)集成,而无需改变其展开策略或优化器。在五个任务族和1.5B至8B的模型中,DARS在相同预算和工具(ALFWorld)下比GiGPO训练的成功率提高了最多10个百分点,提升了WebShop任务得分和Search-R1 QA准确率,通过Python解释器补充了AEPO在AIME24/25上的基于熵的训练,并在1.7B和4B的受控无工具推理比较中超过了OmniOPD。消融实验表明,步骤级信用、依赖衰减和图拓扑各自都有贡献。在ALFWorld上,一个蒸馏的8B标注器匹配了API标注器,使DARS能够在没有前沿评判的情况下高效运行。代码可在该https URL获取。
英文摘要
When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/failure reward, every step in a failed episode has zero total future reward, even when it made progress. We propose Dependency-Aware Reward Shaping (DARS), which represents task progress as predicates linked by prerequisite relations and assigns step-level credit over the dependency graph. An annotator marks which predicates each step verifies, invalidates, or repairs. Verified predicates are discounted according to graph distance from the nearest broken prerequisite, while independent predicates are unaffected. Repairs update these weights based on any errors that remain; invalidated predicates need re-verification to regain credit. A fixed potential converts these annotations into signed per-step rewards. A common reward and annotation interface allows DARS to integrate with a range of reasoning and agentic training methods, such as GiGPO and ARPO/AEPO, without changing their rollout strategies or optimizers. Across five task families and models from 1.5B to 8B, DARS improves success by up to 10 points over GiGPO trained with the same budget and harness (ALFWorld), raises the WebShop task score and Search-R1 QA accuracy, complements AEPO's entropy-based training on AIME24/25 with a Python interpreter, and exceeds OmniOPD in controlled tool-free reasoning comparisons at 1.7B and 4B. Ablations show that step-level credit, dependency attenuation, and graph topology each contribute. On ALFWorld, a distilled 8B annotator matches the API annotator, enabling DARS to run efficiently without a frontier judge. Code is available at https://github.com/JianhuiWei7/DARS.