发表机构
Shanghai Jiao Tong University; University of Cambridge; Tsinghua University; Alibaba Token Hub, Alibaba Group; Nankai University; Chinese University of Hong Kong; SII(上海交通大学; 剑桥大学; 清华大学; 阿里巴巴集团阿里通义实验室; 南开大学; 香港中文大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对CTC训练内存随词表线性增长的问题,提出剪枝CTC,将对齐计算限制在目标词元子集且与全词表等价,并发展出LLM-CTC,在保持准确率的同时大幅降低内存和加速识别。
AI 中文摘要
连接时序分类(CTC)天然支持具有话语级监督的离线与流式语音识别,但传统实现在内存中物化逐帧-词表激活,使得使用原生大语言模型词表的CTC训练内存消耗过高。一个关键观察是,每个有效的CTC对齐仅使用目标词元和空白,且它们在批次中的并集通常构成完整词表的一小部分子集。我们提出剪枝CTC,将对齐计算限制在该子集内,同时保留全词表归一化。我们证明这种词表缩减在损失和梯度上与全词表CTC完全等价。头部和损失激活内存不再随词表大小线性增长。我们进一步应用有限波束对齐剪枝。基于剪枝CTC,我们开发了LLM-CTC,它适配预训练大语言模型用于非自回归语音识别,同时保留因果注意力和原生词表,并将其扩展到有界历史流式识别,避免块级语音-文本对齐。实验表明,使用Zipformer-M编码器和180K词表,剪枝CTC将全步内存减少5.1倍,仅增加17%的步时间开销。在三个语料库上,其准确率与标准CTC相当。在GigaSpeech上,跨六个Qwen3模型规模(从0.6B到32B),LLM-CTC相对LLM-CE的WER保持在7%以内,识别速度快7到10倍;当微调Qwen3-ASR用于有界历史流式识别时,LLM-CTC在测试集上相对匹配的离线模型的WER保持在3%以内。这些结果共同确立了剪枝CTC作为原生词表大语言模型语音识别在离线和流式场景中的可扩展序列目标。
英文摘要
Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1$\times$ with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10$\times$ faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.