发表机构
Meituan LongCat Interaction; Shanghai Jiao Tong University; Fudan University; Peking University; Nanjing University; University of Chinese Academy of Sciences; Jilin University; University of Science and Technology of China; The Hong Kong Polytechnic University(美团龙猫交互; 上海交通大学; 复旦大学; 北京大学; 南京大学; 中国科学院大学; 吉林大学; 中国科学技术大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对策略内蒸馏的学生轨迹偏差问题,提出FutureBridge-OPD方法,在多任务基准上显著优于现有方法,代码公开可用。
AI 中文摘要
策略内蒸馏(OPD)为学生访问的状态提供教师监督,可减少训练与推理间的分布差距。但在多轮智能体任务中,学生偏差会随时间累积,逐渐使轨迹偏离教师指导仍有效的状态。我们的定量分析进一步表明,高分歧状态为教师指导提供了有前景的机会,但确定此类指导是否有益需考察其对后续学生轨迹的影响。我们提出FutureBridge-OPD(FTB),该方法在高分歧状态下执行一段短期教师桥接,并利用由此产生的学生续段评估该桥接是否会增加相对于教师的正蒸馏信号密度。在ALFWorld、WebShop和ScienceWorld上,以Qwen3-32B为教师、Qwen3-1.7B为学生的主要设置下,FTB的性能分别优于普通OPD和TCOD,平均提升16.6和7.6个百分点,且在不同学生规模和教师设置下均保持有效。我们的代码可在此URL获取。
英文摘要
On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn agentic tasks, student deviations may accumulate over time, gradually moving the trajectory away from states where teacher guidance remains effective. Our quantitative analysis further shows that high-disagreement states offer promising opportunities for teacher guidance, but determining whether such guidance is beneficial requires examining its effect on subsequent student trajectories. We propose FutureBridge-OPD (FTB), which executes a short teacher bridge at a high disagreement state and uses the resulting student continuation to assess whether the bridge increases the density of positive distillation signals relative to the teacher. On ALFWorld, WebShop, and ScienceWorld, under the main Qwen3-32B teacher to Qwen3-1.7B student setting, FTB outperforms vanilla OPD and TCOD by an average of 16.6 and 7.6 points, respectively, and remains effective across student scales and teacher settings. Our code is publicly available at https://github.com/ChenChiShui/FutureBridge-OPD.
Comments15 pages, 5 figures