发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出通过潜在空间蒸馏压缩流式神经音频分词器,以预量化潜在为监督目标,在2.8倍压缩下保持低WER偏差,性能优于同等规模独立训练的分词器。
AI 中文摘要
苹果设备上的全系统听写完全在设备端运行,其转录的语音通过分词器(tokenizer)传递给基础模型:该分词器是将短波形窗口映射为语言模型可读表示的编码器。由于该模型在指令跟随剪枝(Instruction-Following Pruning)下稀疏激活,其专家(expert)子集在任意时刻仅占用DRAM的一小部分,因此始终在线的分词器会争夺相同内存,且其参数数量直接影响功耗和延迟。本研究探索如何通过蒸馏压缩此类分词器,将监督目标既非离散token也非输出分布,而是模型实际使用的预量化潜在(pre-quantizer latent)——即两个token接口共享的最后表示。我们仅训练学生编码器(student encoder)以平方误差损失回归教师模型(teacher)的逐帧潜在,通过单个仿射层吸收师生宽度不匹配。由于目标位于量化器和语言模型桥之前,该方法适用于我们支持的两种token接口,且可应用于单独预训练的分词器及与语言模型联合训练的分词器。在2.8倍压缩下,蒸馏得到的学生模型在6组师生配对中的5组上,相对于教师的词错误率(WER)偏差不超过1.9%,且无需任何微调,相比同等规模的独立训练分词器,其相对性能提升3.9%。
英文摘要
System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes - the last representation the two token interfaces share. We train only the student encoder to regress the teacher's per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher-student width mismatch. Because the target precedes both the quantizer and the language-model bridge, one recipe covers both token interfaces we support, and applies both to a tokenizer pretrained alone and to one jointly trained with a language model. At 2.8x compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher-student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative.
Comments15 pages