arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于前缀重放的多轮在线策略蒸馏

Multi-Turn On-Policy Distillation with Prefix Replay

Baohao Liao, Hanze Dong, Christof Monz, Xinxing Xu, Li Dong, Furu Wei

arXiv 2607.04763首次发表:更新:

发表机构

Microsoft Research; University of Amsterdam(微软研究院; 阿姆斯特丹大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究用于智能体任务的在线策略蒸馏,提出重放前缀在线策略蒸馏方法,复用预收集教师轨迹作重放前缀,解决多轮在线策略蒸馏的前缀陷阱问题,提高效率且保持或提升精度。

AI 中文摘要

我们研究用于智能体任务的在线策略蒸馏(OPD),其中大语言模型智能体与环境进行多轮交互,学生智能体在这些多轮交互历史中模仿教师。完全在线的OPD成本高昂,我们提出重放前缀在线策略蒸馏(ReOPD),它重用预收集的教师轨迹作为重放前缀,解决了多轮OPD的前缀陷阱问题,提高了效率。

英文摘要

We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because each update requires fresh student rollouts through the environment and teacher queries at visited histories. We propose Replayed-Prefix On-Policy Distillation (ReOPD), an off-environment alternative that reuses pre-collected teacher trajectories as replayed prefixes: the student acts at selected steps, while the teacher provides dense per-step supervision without executing new environment interactions. We show that multi-turn OPD introduces a prefix trap: making histories more student-on-policy improves relevance to the student, but can query the teacher on histories where its target is unreliable. This creates a two-sided distribution shift between student occupancy and teacher reliability. ReOPD addresses this by treating multi-turn OPD as a reliability-aware prefix distribution design and implements it with a simple step-decaying sampling schedule that emphasizes early, lower-shift prefixes. Across mathematical reasoning with Python and search environments over multiple teacher and student model scales, ReOPD preserves or improves OPD-level accuracy, uses zero tool calls during student training, and is at least 4$\times$ faster per rollout than OPD. ReOPD therefore turns expensive agent-environment interaction into a reusable offline resource, enabling scalable distillation across tools, tasks, and environments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑