发表机构
DeepCybo; Southern University of Science and Technology; Hong Kong Polytechnic University; Zhongguancun Academy; Nanyang Technological University(DeepCybo; 南方科技大学; 香港理工大学; 中关村学院; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出ActFirst-OPD框架,通过参考条件逆动力学先行动后推理,解耦环境交互与响应生成,在多个基准上实现2.3至4.9倍训练加速并保持或提升任务成功率。
AI 中文摘要
在线策略蒸馏(OPD)利用密集的教师监督,在学生生成的响应上训练多轮语言智能体。然而,标准的“先思考后行动”的轨迹生成需要在每个简短行动之前进行冗长的推理,从而延迟了环境转换和经验收集。直接生成行动可以减少这种延迟,但可能降低轨迹质量。为解决这一问题,我们提出了ActFirst-OPD,一种“先行动,后推理”的训练框架,将环境交互与完整响应生成解耦。学生通过参考条件逆动力学,利用其当前交互上下文和参考的下一观测来推断并执行行动,当产生的转换偏离参考轨迹时,则切换到自主的下一行动预测。从收集的交互上下文中,学生异步生成完整的“先思考后行动”响应,用于词级别的教师监督。在0.6B、1.7B和4B参数的Qwen3学生模型上的实验表明,与Vanilla OPD相比,ActFirst-OPD在ALFWorld、WebShop和ScienceWorld上分别实现了平均2.3倍、1.8倍和4.9倍的墙钟训练加速。在九个基准-模型设置中的八个上,它在平均任务成功率方面达到或超过了所有对比的OPD基线。这些结果表明,在多轮智能体蒸馏中,推理不必阻碍行动。
英文摘要
On-policy distillation (OPD) trains multi-turn language agents with dense teacher supervision on student-generated responses. However, standard think-then-act rollouts require lengthy reasoning before each short action, delaying environment transitions and experience collection. Generating actions directly reduces this delay but can degrade rollout quality. To address this, we propose ActFirst-OPD, an act-first, reason-later training framework that decouples environment interaction from full-response generation. The student infers and executes actions through reference-conditioned inverse dynamics using its current interaction context and a reference next observation, and switches to autonomous next-action prediction when the resulting transition deviates from the reference trajectory. From the collected interaction contexts, the student asynchronously generates full think-then-act responses for token-level teacher supervision. Experiments across 0.6B-, 1.7B-, and 4B-parameter Qwen3 students show that ActFirst-OPD achieves average wall-clock training speedups of $2.3\times$ on ALFWorld, $1.8\times$ on WebShop, and $4.9\times$ on ScienceWorld over Vanilla OPD. It matches or exceeds all compared OPD baselines in mean task success rate across eight of nine benchmark-model settings. These results demonstrate that reasoning need not block acting during multi-turn agent distillation.
Comments31 pages, 8 figures