arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SCOUT:面向稀疏奖励强化学习的每上下文重置课程

SCOUT: Per-Context Reset Curricula for Sparse-Reward Reinforcement Learning

Siddharth Aphale, Ayushman Singh

arXiv 2607.26417首次发表:更新:

AI 中文总结

SCOUT是为每个上下文提供独立重置课程的在线控制器,在六类任务中解决了全局进度无法适配学习差异的问题,无需组标签即可提升稀疏奖励强化学习效果。

AI 中文摘要

稀疏奖励强化学习常因无辅助评估起点的轨迹很少到达任务后期而失效。重置课程通过将部分训练轨迹从称为“支架(scaffolds)”的更简单中间状态启动来解决这一问题。这类课程面临两个决策:支架访问(获取有信息的起点)和支架分配(决定移除辅助的速度)。多数现有课程采用单一共享进度移除辅助,当任务实例或上下文学习速率不同时会失效。我们提出SCOUT,一种在线、与学习者无关的重置控制器,为每个上下文提供独立课程。仅使用轨迹成功与否的二元信息,SCOUT在持续成功后移除辅助,失败后恢复辅助,进展停滞时谨慎测试更难的起点,且不改变奖励、优化器或学习者。计数分析表明,当上下文需要不同量的辅助练习时,同步全局进度可能不足。在六个导航和操纵环境中,支架访问提升了学习效果,使三个环境中无辅助训练在给定预算内失败的任务得以成功。在构建的进度冲突场景中,每个测试的全局进度都有一组任务未解决,而SCOUT解决了两组任务。平均成功率可能掩盖这种失败,因此我们还报告了成功率最低的组。当学习差异遵循已知组时,组级进度有效,但当差异出现在同一组内时会失效。SCOUT无需组标签,在两种情况下均保持稳定性能。重置课程应在学习进度存在差异的尺度上移除辅助。

英文摘要

Sparse-reward reinforcement learning often fails because rollouts from the unassisted evaluation start rarely reach later task stages. Reset curricula address this by starting some training rollouts from easier intermediate states, called scaffolds. Such a curriculum faces two decisions: scaffold access, obtaining informative starts, and scaffold allocation, deciding how quickly that assistance is removed. Most prior curricula pace removal on one shared schedule, which can fail when task instances, or contexts, learn at different rates. We introduce SCOUT, an online, learner-agnostic reset controller that gives every context its own curriculum. Using only binary rollout success, SCOUT removes assistance after sustained success, restores it after failure, and cautiously tests a harder start when progress stalls, without changing the reward, optimizer, or learner. A counting construction shows that synchronized global pacing can be insufficient when contexts need conflicting amounts of assisted practice. Across six navigation and manipulation settings, scaffold access improves learning and enables success in three where unassisted training fails within the reported budget. In a constructed pacing conflict, each tested global schedule leaves one group unsolved, while SCOUT solves both. Average success can conceal this failure, so we also report the least successful group. Group-level pacing works when learning differences follow known groups but can fail when they occur within one group. SCOUT needs no group labels and remains consistently strong in both cases. A reset curriculum should remove assistance at the scale where learning progress differs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑