arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27035cs.AIcs.LG

基于子任务分解的强化学习

Reinforcement Learning with Decomposed Subtasks

Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich

首次发表
浏览论文内容

中文总结 AI 辅助

针对GRPO将多轮轨迹压缩为单一奖励导致技能归因损失的问题,提出基于子任务分解的强化学习(RLDS),通过子任务分解优势估计(SDAE)按子任务分配信用,在异质性任务上显著提升性能。

中文摘要 AI 辅助

组相对策略优化(GRPO)及相关用于训练语言模型智能体的策略梯度方法,在将整个多轮轨迹输入策略更新之前,会将其压缩为单个标量轨迹奖励。当任务由不同技能组合而成,尤其是在稀疏和延迟的环境反馈下,这种压缩是有损的:优化器必须隐式推断是哪种能力导致了该结果,以及这应如何改变行为。我们认为正确的原语不是更好的标量,而是分解:轨迹奖励应在进入策略更新之前按子任务进行拆分。我们引入了基于子任务分解的强化学习(RLDS),其核心是子任务分解优势估计(SDAE):一种替代标量GRPO优势的方法,它根据固定分类法将轨迹奖励拆分为每个子任务的份额,计算每个子任务的组相对优势,并通过按重要性加权每个子任务的优势来分配每个令牌的信用,将其集中在反映该子任务执行具有重要意义的步骤周围。我们在四个智能体基准上进行了评估:FrozenLake(稀疏网格导航)、HotpotQA(多跳问答,一个检索工具)、ScienceWorld(长视界具身科学)和DeepResearch(长形式研究,四个工具,复合评分奖励)。训练期间发出的异质性诊断显示了分解在何处发挥作用——收益随子任务异质性而扩展,在高异质性任务ScienceWorld(+11.5分,配对自助法95%置信区间[+9.8,+13.3])和FrozenLake(+9.8分,[+7.0,+12.8])上最大,而在HotpotQA和DeepResearch上处于噪声范围内,诊断预测这些任务几乎没有可恢复的收益。在RLDS下,ScienceWorld也比标量GRPO更具计算效率(每步墙钟时间-10.9%),因为长轨迹摊销了固定的反思和评分开销。

英文摘要

Group Relative Policy Optimization (GRPO) and related policy-gradient methods for training language model agents collapse an entire multi-turn rollout into a single scalar trajectory reward before it enters the policy update. When the task composes distinct skills, especially under sparse and delayed environmental feedback, this collapsing is lossy: the optimizer must implicitly infer which competency drove the outcome and how that should change behavior. We argue the right primitive is not a better scalar but a decomposition: trajectory reward should be split along subtasks before it enters the policy update. We introduce Reinforcement Learning with Decomposed Subtasks (RLDS), whose core is Subtask-Decomposed Advantage Estimation (SDAE): a replacement for the scalar GRPO advantage that splits trajectory reward into per-subtask shares on a fixed taxonomy, computes a group-relative advantage per subtask, and distributes per-token credit by weighting each subtask's advantage by its importance, concentrating it around the step where a reflection marks that subtask's execution as consequential. We evaluate on four agentic benchmarks: FrozenLake (sparse grid navigation), HotpotQA (multi-hop QA, one retrieval tool), ScienceWorld (long-horizon embodied science), and DeepResearch (long-form research, four tools, composite rubric reward). Heterogeneity diagnostics emitted during training show where decomposition pays off - gains scale with subtask heterogeneity, largest on the high-heterogeneity tasks ScienceWorld (+11.5 points, paired-bootstrap 95% CI [+9.8, +13.3]) and FrozenLake (+9.8 points, [+7.0, +12.8]), and within noise on HotpotQA and DeepResearch, where the diagnostics predicted little to recover. ScienceWorld is also more compute-efficient under RLDS than scalar GRPO (-10.9% wall-clock per step), as long rollouts amortize the fixed reflect-and-grade overhead.

发表机构

  • Upwork

机构由 AI 辅助整理,请以论文原文为准。

↑