发表机构
The Chinese University of Hong Kong; IDEA Research; Emdoor Research Institute(香港中文大学; IDEA研究院; 亿道研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对在线策略蒸馏中执行器切换的问题,提出带滞后性的差异感知切换算法 DASH-OPD,经 ALFWorld 验证,其性能优于所有基线,训练与部署效率更优,相关代码等后续将发布。
AI 中文摘要
在线策略蒸馏(On-Policy Distillation, OPD)通过利用自身的 rollout 训练学生模型以减少暴露偏差。但在多轮智能体场景中,学生早期的错误会使轨迹偏离教师的熟悉领域。现有课程学习方法会根据训练进度调整教师支持的使用量,却无法判断何时需要教师支持。针对这一问题,本文提出 DASH-OPD,即面向 OPD 的带滞后性的差异感知切换算法,这是一种新型智能体式 OPD 方法,可自适应双向切换执行器。每一轮中,DASH-OPD 会计算两个执行器在动作 token 上的平均对数概率比作为差异;学生轮次的学生到教师的比值形成漂移信号,教师轮次的教师到学生的比值形成恢复信号。这些信号经多轮归一化后累加为漂移和恢复证据,当证据超过对应切换阈值时,DASH-OPD 会切换执行器。这种多轮累加使切换具有滞后性,可避免瞬时波动导致的高频切换。在 ALFWorld 上,DASH-OPD 优于所有基线,且展现出更优的训练和部署效率。本文为进行中的工作,代码、训练日志及模型检查点将后续发布。
英文摘要
While on-policy distillation (OPD) reduces exposure bias by training student language models on their own rollouts, early student errors in long-horizon agentic scenarios can lead to contexts unfamiliar to the teacher. To improve trajectory quality, recent work on agentic OPD introduces teacher intervention into training rollouts by switching the executor between the student and the teacher. However, existing methods determine how much teacher intervention is needed---but not when. To address this limitation, we propose DASH-OPD (Discrepancy-Aware Switching with Hysteresis for OPD), the first agentic OPD method to perform adaptive, bidirectional executor switching. At each turn, DASH-OPD measures teacher--student discrepancy using a mean log-likelihood ratio over action tokens. High student-to-teacher ratios on student turns serve as drift signals, while low teacher-to-student ratios on teacher turns serve as recovery signals. These signals are accumulated over multiple turns to form drift and recovery evidence, respectively. DASH-OPD switches executors when either type of evidence exceeds its corresponding switching threshold, introducing hysteresis that prevents rapid switching triggered by transient discrepancy fluctuations. Across three benchmarks and two student model scales, DASH-OPD outperforms five baselines in all 14 task-performance comparisons, while requiring the fewest interaction turns in nine of ten efficiency comparisons. Code, trained models, and training logs are available at https://github.com/Lucian1115/DASH-OPD