SpikeOPD:自激语言模型的稳定在线策略蒸馏方法
SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models
浏览论文内容
中文总结 AI 辅助
针对自回归脉冲语言模型的前缀不匹配问题,提出SpikeOPD框架,结合教师校正、策略锚定与分层脉冲正则化,提升不同规模模型的准确率并保留稀疏计算特性。
中文摘要 AI 辅助
脉冲神经网络(SNN)通过稀疏编码和事件驱动计算为高效语言建模提供了途径,但从头训练高性能脉冲语言模型仍颇具挑战。一种实用替代方案是通过知识蒸馏(KD)实现人工神经网络(ANN)到SNN的迁移,即利用预训练的ANN教师模型监督SNN学生模型。现有迁移方法在固定语料前缀上进行蒸馏,而自回归推理依赖自身生成的前缀,导致前缀来源不匹配,具体表现为与ANN教师的输出策略不匹配,以及自身生成前缀与匹配语料前缀间的内部脉冲动力学漂移。在线策略蒸馏(OPD)通过在自身生成的前缀上持续引入教师监督,自然缓解了上述两种问题。我们通过受控压力测试评估仅使用教师的全KL散度变体(Vanilla OPD),发现其可能出现延迟的生成反馈崩溃,表明仅在线策略覆盖无法确保稳定适配。基于此,我们提出SpikeOPD,一种面向自回归SNN的稳定在线策略蒸馏框架,可在学习自身生成前缀的同时保持生成稳定性。该框架采用全KL散度教师校正以减少输出策略不匹配,同时采用匹配前缀策略锚定以约束策略在相同前缀上偏离冻结参考SNN的程度,还引入分层脉冲正则化以限制在线策略适配期间的放电率偏差。在三个模型规模下,SpikeOPD分别在0.125B、0.35B和1.3B规模时,较对应KD SNN的平均准确率提升0.8、1.7和2.9个百分点,同时保留其稀疏计算特性。
英文摘要
Spiking neural networks (SNNs) offer a path to energy-efficient language modeling through sparse encoding and event-driven computation, but training capable spiking language models from scratch remains difficult. A practical alternative is ANN-to-SNN migration through knowledge distillation (KD), where a pretrained artificial neural network (ANN) teacher supervises an SNN student. Existing migration approaches distill on fixed corpus prefixes, whereas autoregressive inference conditions on self-generated prefixes, creating prefix-source mismatch. It manifests as output-policy mismatch with the ANN teacher and internal spiking-dynamics drift between self-generated and matched corpus prefixes. On-policy distillation (OPD) offers a natural way to mitigate both manifestations by continuing teacher supervision on self-generated prefixes. We evaluate a teacher-only full-KL variant, Vanilla OPD, via a controlled stress test and observe it may suffer from delayed rollout-feedback collapse. This result shows that on-policy coverage alone does not ensure stable adaptation. Motivated by these findings, we propose SpikeOPD, a stable on-policy distillation framework for autoregressive SNNs that learns from self-generated prefixes while maintaining rollout stability. It applies full-KL teacher correction to reduce output-policy mismatch, while matched-prefix policy anchoring constrains policy departure from the frozen reference SNN on the same prefixes. Layerwise spike regularization further limits firing-rate deviations during on-policy adaptation. Across three model scales, SpikeOPD improves average accuracy over the corresponding KD SNNs by 0.8, 1.7, and 2.9 points at 0.125B, 0.35B, and 1.3B, respectively, while preserving their sparse-compute profiles.
发表机构
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Agency for Science, Technology and Research (A*STAR), Singapore(新加坡科技研究局)
- Nanyang Technological University, Singapore(新加坡南洋理工大学)
- Peking University(北京大学)
- National University of Singapore(新加坡国立大学)
- Rice University(莱斯大学)
机构由 AI 辅助整理,请以论文原文为准。