当特权指导出现错位:面向多轮智能体的状态匹配路由与语境化自蒸馏
When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
浏览论文内容
中文总结 AI 辅助
针对多轮智能体特权指导的状态-参考不匹配问题,提出SMRC-SD方法,通过状态匹配路由与语境化自蒸馏提升了ALFWorld和WebShop任务的成功率。
中文摘要 AI 辅助
特权在线策略蒸馏通过允许同步教师在每一轮利用仅训练时可用的参考(如成功轨迹)对智能体的响应重新评分,为多轮智能体提供密集监督。然而在交互环境中,智能体的前置动作会持续改变执行状态。当智能体采取不同动作或以不同顺序完成子目标时,其生成的轨迹可能会到达参考未覆盖的状态,导致参考无法为实际到达的状态提供可靠指导。无差别应用特权蒸馏会造成状态-参考不匹配,这一问题促使核心目标:提供与智能体当前执行状态兼容的特权参考指导。我们提出状态匹配路由与语境化自蒸馏(SMRC-SD),明确确定特权轨迹应在何时、以何种方式指导在线策略智能体。每一轮中,SMRC-SD会验证智能体当前执行状态是否与参考轨迹上的支持状态匹配,仅在匹配状态下应用蒸馏,过滤掉参考缺乏局部兼容指导的轮次。对于每个匹配状态,SMRC-SD还会从成功轨迹中构建状态条件教师语境,将监督建立在实际到达的状态之上。在ALFWorld和WebShop数据集上,SMRC-SD始终优于无条件的成功全路径蒸馏。使用Qwen3-1.7B模型时,其在ALFWorld上的任务成功率从0.746提升至0.865,在WebShop上从0.574提升至0.693。受控路由与语境消融实验表明,选择局部支持轮次和构建状态兼容教师语境是性能提升的两个关键因素。代码可在该https网址获取。
英文摘要
Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access to training-only references, such as successful trajectories. In interactive environments, however, the student's preceding actions continually change the execution state. As the student takes different actions or completes subgoals in a different order, its rollout may reach states not covered by the reference, making the reference an unreliable source of guidance for the state actually reached. Applying privileged distillation indiscriminately therefore creates state--reference mismatch. This mismatch motivates a central objective: providing privileged reference guidance that remains compatible with the student's current execution state. We introduce State-Matched Routing and Contextualized Self-Distillation (SMRC-SD), which explicitly determines when and how a privileged trajectory should guide an on-policy student. At each turn, SMRC-SD verifies whether the student's current execution state matches a supported state along the reference trajectory. Distillation is applied only at matched states, filtering out turns for which the reference lacks locally compatible guidance. For each matched state, SMRC-SD further constructs state-conditioned teacher context from the successful trajectory, grounding supervision in the state actually reached. Across ALFWorld and WebShop, SMRC-SD consistently outperforms unconditional successful full-path distillation. With Qwen3-1.7B, it improves task success from $0.746$ to $0.865$ on ALFWorld and from $0.574$ to $0.693$ on WebShop. Controlled routing and context ablations support both selecting locally supported turns and constructing state-compatible teacher context as contributors to these gains. Code is available at https://github.com/liujunzhuo/SMRC-SD.