arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35058cs.LGcs.AI

TIDE:通过信息蒸馏与探索实现智能体强化学习的师生过渡

TIDE: Teacher-Student Transition via Informative Distillation and Exploration for Agentic RL

  • Harbin Institute of Technology(哈尔滨工业大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Yibin Huang, Xinming Xu, Conghui Zhu

AI总结:

针对多轮智能体强化学习中教师蒸馏与奖励优化固定混合的局限,提出TIDE方法,通过全局分歧趋势调度OPD到RL的交接、局部联合分配两种信号,实现自适应协调,提升训练效果。

AI中文摘要:

有效的多轮智能体需要协调信息收集、行动和反馈的交互策略,以应对长时程任务。GRPO是一种用于训练这些智能体的强化学习算法,但稀疏的轨迹级奖励限制了小模型的早期探索。近期方法通过从更强的教师模型进行在线策略蒸馏(OPD)来增强强化学习。然而,固定的混合比例假设教师指导和奖励优化在整个训练过程中以及各交互轮次之间应保持恒定的相对作用。这一假设可能在两个尺度上失效。全局上,随着训练推进,维持较强的蒸馏压力可能约束模型超越教师能力。局部上,师生分歧可以识别学生偏离教师之处,但无法判断这种偏离是得到更好结果支持的探索,还是低质量的策略漂移。我们的方法论洞见是,教师指导和奖励优化应在训练过程中动态重新平衡,并在各轮次间联合分配。我们将这一洞见实例化为TIDE。全局上,TIDE将实测的分歧趋势作为实用的调度信号,在分歧减少变慢但仍为正值时推进从OPD到RL的交接,并逐步增加RL的相对权重。局部上,TIDE联合调节教师引导和奖励驱动的更新:相对动作价值和分歧优先考虑OPD信号,而相对动作价值提供RL优势,归一化分歧在轮次间重新加权该优势。与全局交接相结合,TIDE在训练早期分配更强的教师指导,并在训练后期给予奖励驱动更新更大的相对权重。跨多个基准、学生规模和受控消融实验的结果支持了TIDE自适应OPD-RL协调的有效性。

英文摘要:

Effective multi-turn agents require interaction strategies that coordinate information gathering, actions, and feedback over long horizons. GRPO is a reinforcement learning algorithm used to train these agents, but sparse trajectory-level rewards limit early exploration in small models. Recent methods augment RL with on-policy distillation (OPD) from a stronger teacher. However, a fixed mixture assumes that teacher guidance and reward optimization should retain a constant relative role throughout training and across interaction turns. This assumption can fail at two scales. Globally, as training progresses, maintaining strong distillation pressure can constrain the model from moving beyond the teacher's capabilities. Locally, teacher--student disagreement identifies where the student departs from the teacher, but cannot tell whether that departure is exploration supported by better outcomes or low-quality policy drift. Our methodological insight is that teacher guidance and reward optimization should be dynamically rebalanced over training and jointly allocated across turns. We instantiate this insight in \tide. Globally, \tide uses the measured disagreement trend as a practical schedule signal, advancing an OPD-to-RL handoff when discrepancy reduction becomes slow but remains positive and progressively increasing the relative weight of RL. Locally, \tide jointly modulates teacher-guided and reward-driven updates: relative action value and disagreement prioritize the OPD signal, whereas relative action value supplies the RL advantage and normalized disagreement reweights it across turns. Coupled with the global handoff, \tide allocates stronger teacher guidance early and gives reward-driven updates greater relative weight later in training. Experiments across multiple benchmarks, student scales, and controlled ablations support the effectiveness of TIDE's adaptive OPD--RL coordination.

↑