CLOOPD:在策略蒸馏中闭合学习者循环
CLOOPD: Closing the Learner Loop in On-Policy Distillation
浏览论文内容
中文总结 AI 辅助
CLOOPD提出闭环框架,分离教师信号获取与学生实现,通过自适应α路径点和多遍策略,在减少令牌和GPU消耗的同时提升蒸馏性能。
中文摘要 AI 辅助
在策略蒸馏(OPD)中,每一批新数据都要付出双倍代价:学生生成轨迹,更强的教师对其评分。现有方法改进了哪些轨迹被评分以及教师信号如何构建,但通常仅通过一次演员更新来消耗这些信号。我们引入了CLOOPD,一个闭环框架,将教师信号获取与学生端实现分离。CLOOPD在KL包络内选择一个自适应的$\alpha$路径点,冻结已评分的批次及其优势值,在每次演员传递后重新前向传播学生,测量实现程度,并在单独的令牌预算下分配演员工作。该框架包括确定性的两遍和三遍策略、按令牌计价的CLOOPD-TPMR以及一个预算匹配的对照组。在8-H20节点上的六次300步运行中,每个CLOOPD策略在可比的教师令牌规模下都优于单遍TOP-D锚点:宏观准确率从15.41提升至17.78(CLOOPD-Fixed2)和19.36(CLOOPD-Fixed3)。在第100步时,CLOOPD-Fixed3达到15.35,几乎与第300步的TOP-D相当,同时使用的教师评分令牌减少67.2%,GPU小时数减少28.0%。早期的8-A100消融实验表明,自适应$\alpha$消除了观察到的信任包络违规;第三遍增加了余量。这些结果使CLOOPD成为一个用于预算学生从教师评分令牌中学习充分程度的框架。
英文摘要
On-policy distillation (OPD) pays twice for each fresh batch: the student generates trajectories and a stronger teacher scores them. Existing methods improve which trajectories are scored and how the teacher signal is constructed, but usually consume it with one actor update. We introduce CLOOPD, a closed-loop framework separating teacher-signal acquisition from student-side realization. CLOOPD selects an adaptive $α$ waypoint inside a KL envelope, freezes the scored batch and its advantages, re-forwards the student after each actor pass, measures realization, and allocates actor work under a separate token budget. The framework includes deterministic two- and three-pass policies, token-priced CLOOPD-TPMR, and a budget-matched control. Across six 300-step runs on an 8-H20 node, every CLOOPD policy improves the one-pass TOP-D anchor at comparable teacher-token scale: macro accuracy rises from 15.41 to 17.78 with CLOOPD-Fixed2 and 19.36 with CLOOPD-Fixed3. At step 100, CLOOPD-Fixed3 reaches 15.35, nearly matching TOP-D at step 300 while using 67.2% fewer teacher-scored tokens and 28.0% fewer GPU-hours. Earlier 8-A100 ablations show adaptive $α$ eliminates observed trust-envelope violations; a third pass adds headroom. These results position CLOOPD as a framework for budgeting how fully students learn from teacher-scored tokens.
发表机构
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。