发表机构
School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Progressive²是结合渐进式增强教师与渐进式缩小学生的知识蒸馏方法,可缓解服务器-客户端能力差距导致的蒸馏性能下降,实现模型压缩并提升准确率与训练效率。
AI 中文摘要
知识蒸馏(Knowledge Distillation,KD)是一种广泛应用的技术,用于将大型模型(教师模型)的知识迁移至较小模型(学生模型)。凭借其灵活性和广泛适用性,KD已被大量应用于服务器端模型的压缩,以满足客户端用户的服务质量(Quality of Service,QoS)要求。尽管该领域已取得显著进展,但当服务器能力与客户端需求存在较大差距时,蒸馏性能会大幅下降。为缓解这一问题,我们提出一种名为Progressive²的新型蒸馏方法,该方法通过结合渐进式增强的教师模型与渐进式缩小的学生模型来实现。在教师模型侧,我们并非同时使用所有层,而是遵循从原始到丰富的语义演进规律,逐步选择额外层进行蒸馏,从而建立系统化的学习课程。此外,我们设计了一种教师侧多特征融合适配器,以提升教师模型的训练稳定性,该设计在理论上得到了Lipschitz连续性框架的支撑。在学生模型侧,我们并非直接训练微型模型,而是逐步减小网络规模,以促进与教师模型的迭代协同演化。Progressive²是一种灵活的框架:教师模型的渐进策略可单独部署,以在准确率与训练效率间实现最优平衡;而教师与学生模型的联合集成则能进一步提升整体性能。
英文摘要
Knowledge distillation (KD) is a widely utilized technique for transferring knowledge from a large model (the teacher) to a smaller model (the student). Owing to its flexibility and broad applicability, KD has been extensively applied in the compression of server-side models to meet the Quality of Service (QoS) requirements of client users. Despite significant advancements, the performance of distillation is substantially compromised when a large disparity exists between the capabilities of the server and the requirements of the client. To alleviate this problem, we propose a novel distillation approach, named Progressive$^2$, which operates through the combination of a progressively stronger teacher and a progressively smaller student. On the side of the teacher, rather than involving all layers simultaneously, we progressively select additional layers for distillation following a raw-to-rich semantic progression, establishing a systematic learning curriculum. Furthermore, we design a teacher-side multi-feature fusion adapter for the teacher to improve training stability, which is theoretically supported by the framework of Lipschitz continuity. On the side of the student, rather than directly training a tiny model, we gradually reduce the size of the network to facilitate an iterative co-evolution with the teacher. Progressive$^2$ serves as a flexible framework; the progressive strategy of the teacher can be deployed independently to achieve an optimal balance between accuracy and training efficiency, while the joint integration of the teacher and the student yields further improvements in overall performance.
CommentsManuscript under review at IEEE Transactions on Services Computing