发表机构
Shenzhen Key Laboratory of Internet Information Collaboration; Harbin Institute of Technology, Shenzhen(深圳市互联网信息协同重点实验室; 哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对视觉在线策略蒸馏中在线展开成本高的问题,提出HB-SJD批量推测雅可比展开后端,仅替换学生展开后端,在LlamaGen上验证可降时间且保生成质量。
AI 中文摘要
视觉在线策略蒸馏(OPD)通过学习当前学生模型生成的轨迹,改进紧凑视觉自回归模型的训练。然而,这些在线展开仍以自回归解码方式逐token生成,大幅增加了每一步在线策略训练的成本。推测雅可比解码(SJD)提供了一种替代方案,因其无需辅助草稿模型即可并行处理多个token,但该原始方法专为单序列推理设计。我们提出HB-SJD,一种用于视觉OPD的批量SJD展开后端。HB-SJD允许每个图像根据自身解码进度独立推进,同时不同序列位置的图像仍通过批量模型前向传播进行验证。随着图像完成解码,HB-SJD在全执行与紧凑执行间切换,以降低后续展开轮次的成本。HB-SJD仅替换学生模型的展开后端,教师模型、蒸馏目标及优化流程保持不变。在LlamaGen上的实验表明,HB-SJD在保留蒸馏后学生模型生成质量的同时,大幅降低了展开和端到端训练时间。
英文摘要
Visual on-policy distillation (OPD) improves the training of compact visual autoregressive models by learning from trajectories generated by the current student. However, these online rollouts are still produced token by token with autoregressive decoding, which adds substantial cost to every on-policy training step. Speculative Jacobi Decoding (SJD) provides an alternative because it can process multiple tokens in parallel without an auxiliary draft model, but the original method is designed for single-sequence inference. We introduce HB-SJD, a batched SJD rollout backend for visual OPD. HB-SJD allows each image to advance independently according to its own decoding progress, while images at different sequence positions are still verified in batched model forwards. As images finish, HB-SJD switches between Full and Compact execution to reduce the cost of later rollout rounds. HB-SJD only replaces the student rollout backend and leaves the teacher, distillation objective, and optimization procedure unchanged. Experiments with LlamaGen show that HB-SJD substantially reduces rollout and end-to-end training time while preserving the generation quality of the distilled student.
Comments11 pages,4 figures