具有自提炼奖励塑造的智能体强化学习
Agentic Reinforcement Learning with Self-Distilled Reward Shaping
浏览论文内容
中文总结 AI 辅助
针对智能体强化学习中稀疏奖励无法分配中间决策信用的问题,提出ADRS框架,通过校准教师分数等提升多轮语言智能体在长时域任务的性能,且增益稳定。
中文摘要 AI 辅助
智能体强化学习使大语言模型(LLM)智能体能够通过交互学习,但稀疏的轨迹级奖励仅能揭示成功与否,无法确定哪些中间决策应获得 credit( credit 译为“信用分配”)。仅用于训练的特权技能可提供更密集的监督,方法是让同一个冻结的策略快照在任务匹配的程序技能条件下,对无技能轨迹中的固定 token 进行重新评分。然而,现有方法未在交互步骤间联合校准教师分数、未将教师置信度与实际收益关联,也未将所得信号整合到原生的收益-优势构建中。我们提出了具有自提炼奖励塑造的智能体强化学习(ADRS),这是一种为多轮语言智能体构建与收益关联的 token 级信用的框架。ADRS 以步骤内的特权 token 分数为中心并对其进行归一化,基于组内置信度-收益关联,用与收益关联的教师价值优势(TVA)门对这些分数进行调制,并将门控后的 token 信号整合到原生 RL 信用构建中。这些组件共同确定教师的偏好是什么、该偏好何时与收益相关,以及如何将其纳入原生强化学习信用路径,同时保持 rollout( rollout 译为“轨迹展开”)和推理时无技能。最后,在三个交互式基准上开展的实验表明,ADRS 能持续提升长 horizon( horizon 译为“任务时域”)任务的性能,且增益在不同 RL 主干、数据减少设置、未见过的任务及扩展训练中均能保持。为匿名评审,我们的代码可在以下链接获取:this https URL
英文摘要
Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions deserve credit. Training-only privileged skills can provide denser supervision by allowing the same frozen policy snapshot to rescore fixed tokens from skill-free trajectories while conditioned on task-matched procedural skills. Existing methods, however, do not jointly calibrate teacher scores across interaction steps, relate teacher confidence to realized returns, and integrate the resulting signal into native reward-to-advantage construction. We introduce Agentic Reinforcement Learning with Self-Distilled Reward Shaping (ADRS), a framework for constructing return-associated token-level credit for multi-turn language agents. ADRS centers and normalizes privileged token scores within each step, modulates them with a return-associated Teacher Value Advantage (TVA) gate based on within-group confidence--return association, and incorporates the gated token signal into native RL credit construction. Together, these components determine what the teacher prefers, when that preference is return-relevant, and how it enters the native reinforcement-learning credit path, while keeping rollouts and inference skill-free. Finally, experiments across three interactive benchmarks show that ADRS consistently improves performance on long-horizon tasks, with gains persisting across RL backbones, reduced-data settings, unseen tasks, and extended training. For anonymous review, our code is available at the following the link: https://github.com/gitrxh/ADRS-arxiv