面向智能体在线策略蒸馏的随机教师干预
Stochastic Teacher Intervention for Agentic On-Policy Distillation
浏览论文内容
中文总结 AI 辅助
本文针对多轮智能体在线策略蒸馏中轨迹偏离导致监督不可靠的问题,提出STI-OPD框架,通过随机干预策略与重要性加权反向KL目标提升性能,在多类任务上优于现有基线。
中文摘要 AI 辅助
在线策略蒸馏(OPD)通过对智能体生成的 rollout 进行密集 token 级监督,高效将更强教师模型的能力迁移至学生语言模型,已在数学推理等复杂任务中展现出潜力。然而在多轮智能体任务中,学生的决策会塑造后续观测,导致早期错误在多轮中累积,生成的轨迹会偏离教师的 rollout 分布,使得教师的 token 级监督对 OPD 训练的可靠性降低甚至产生反效果。为解决该问题,本文提出 STI-OPD,一种面向多轮智能体 OPD 的随机教师干预框架。在多轮交互中,STI-OPD 采用由师生策略差异引导的教师干预,用教师生成的动作替换学生提出的动作,以最大化获取可靠监督。本文进一步开发了一种随机干预策略,解决了以往基于阈值或固定调度方法的局限性,该策略用 KL 散度估计策略差异并将其映射为干预概率,通过从该概率中采样是否干预,STI-OPD 自适应平衡教师控制与学生探索。为从生成的混合策略轨迹中学习,本文提出了重要性加权反向 KL 目标,用于修正教师生成响应与学生策略之间的 token 采样不匹配,以保留原始 OPD 目标。在工具集成推理与长程交互任务中,STI-OPD 在所有评估基准和学生模型规模上均优于最强的现有 OPD 基线, ablation 实验进一步表明,差异引导干预和重要性加权均对性能提升有贡献。
英文摘要
On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher intervention guided by teacher-student policy discrepancy to replace the student's proposed action with a teacher-generated one to maximize the acquisition of reliable supervision. We further develop a stochastic intervention strategy, addressing the limitations of previous threshold-based or fixed-schedule approaches, that estimates policy discrepancy using KL divergence and maps it to an intervention probability. By sampling whether to intervene from this probability, STI-OPD adaptively balances teacher control with student exploration. To learn from the resulting mixed-policy trajectories, we introduce an Importance-Weighted Reverse KL objective that corrects the token sampling mismatch between teacher-generated responses and the student policy to preserve the original OPD objective. Across tool-integrated reasoning and long-horizon interaction, STI-OPD outperforms the strongest prior OPD baseline on every evaluated benchmark and student size. Ablations further show that both discrepancy-guided intervention and importance weighting contribute to these gains.
发表机构
- Monash University(莫纳什大学)
- The Hong Kong Polytechnic University(香港理工大学)
- Zhongguancun Laboratory(中关村实验室)
机构由 AI 辅助整理,请以论文原文为准。