arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

引导,然后放手:稀疏奖励智能体强化学习中的间隙自适应教师调度

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

Youling Huang, Tiankuo Xu, Jiaji Liu, Tong Zheng, Shuo Zhou, Shaotong Qi, Junchi Yao, Shiyang Liu, Hao Xu, Pengcheng Xu, Bo Huang, Hongyi Fu, Lin Lin

arXiv 2609.37898首次发表:更新:

发表机构

DUT; XJTU; THU; UCAS; BFSU; SEU; MBZUAI; Kuaishou(大连理工大学; 西安交通大学; 清华大学; 中国科学院大学; 北京外国语大学; 东南大学; 穆罕默德·本·扎耶德人工智能大学; 快手)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对稀疏奖励智能体强化学习的冷启动问题,提出基于师生性能差距自适应的教师调度方法GATS,动态调整蒸馏权重,在教师优于学生时提供指导,差距逆转后撤回,提升平均成功率4.37%-11.87%。

AI 中文摘要

面向长时程智能体的强化学习通常依赖稀疏的基于结果的奖励。这导致了严重的冷启动问题,因为早期策略往往无法解决采样任务,留给学习的有效奖励信号很少。为缓解此问题,我们使用同策略蒸馏(OPD)在学生自身的轨迹上提供词元级指导。我们发现,这种指导的益处取决于教师与学生之间的性能差距。当教师显著优于学生时,蒸馏有助于引导学生度过早期训练阶段,此时结果奖励提供的学习信号很少。然而,随着差距缩小并最终逆转,持续蒸馏变得不那么有益,甚至可能阻碍进一步改进。基于这一观察,我们提出了间隙自适应教师调度(GATS),该方法通过一个权重随师生性能差距自适应的OPD项来增强学生的强化学习目标。具体而言,GATS在学生接近教师参考性能时逐渐减少教师指导,并在达到该参考后撤回指导。这使得GATS能够利用比学生更小的任务训练教师,因为教师指导主要在早期训练时需要。在ALFWorld、WebShop和ScienceWorld上,使用三种Qwen2.5师生配置,GATS在所有三种配置中均取得了比较方法中最高的平均成功率,在匹配的学生轨迹预算下,比仅奖励的GRPO提高了4.37%-11.87%。代码可在以下网址获取:此https URL。

英文摘要

Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate this problem, we use on-policy distillation (OPD) to provide token-level guidance on the student's own rollouts. We find that the benefit of this guidance depends on the performance gap between the teacher and the student. When the teacher substantially outperforms the student, distillation helps guide the student through the early training stage where outcome rewards provide little learning signal. As the gap narrows and eventually reverses, however, continued distillation becomes less beneficial and may hinder further improvement. Motivated by this observation, we propose Gap-Adaptive Teacher Scheduling (GATS), which augments the student's RL objective with an OPD term whose weight adapts to the teacher-student performance gap. Specifically, GATS gradually reduces teacher guidance as the student approaches the teacher's reference performance and withdraws it once that reference is reached. This enables GATS to leverage task-trained teachers smaller than the student, since teacher guidance is primarily needed during early training. Across ALFWorld, WebShop, and ScienceWorld with three Qwen2.5 teacher-student configurations, GATS achieves the highest average success rate among the compared methods in all three configurations, improving over reward-only GRPO by 4.37%-11.87% under matched student rollout budgets. Code is available at https://github.com/Ricardo-H/guide-then-let-go.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑