arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IBM Granite 5.0 TurboCTC ASR 模型设计

Design of the IBM Granite 5.0 TurboCTC ASR Model

Brian Kingsbury, George Saon, Masayuki Suzuki, Hong-Kwang J. Kuo, Takashi Fukuda, Samuel Thomas, Vishal Sunder, Avihu Dekel

arXiv 2609.20104首次发表:更新:

发表机构

IBM Research(IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文介绍 IBM Granite 5.0 TurboCTC,一个4.7亿参数的编码器模型,通过金字塔子采样、块对角注意力和推理优化,在英文短语音识别上达到速度-准确率帕累托最优,速度是竞品两倍。

AI 中文摘要

我们描述了 Granite 5.0 Turbo CTC 的架构、训练方法和推理加速技术。该模型是一个拥有 4.7 亿参数的仅编码器模型,具有出色的速度-准确率权衡。架构采用金字塔式时间子采样,在 Conformer 块内使用步进深度卷积、块对角(分块)自注意力,并基于中间层的预测进行条件化。训练亮点包括仅使用公开可用的数据、新颖地使用 Muon 优化器以及平衡数据采样。推理加速包括用线性层替换 1x1 卷积,并优化 Conformer 块中的注意力计算。综合这些设计,使得该模型在 Open ASR 排行榜的英文短语音识别任务上处于速度-准确率的帕累托前沿,同时其速度是最快竞争者的两倍。该模型可在宽松许可证下使用,并可从该 https URL 下载。

英文摘要

We describe the architecture, training methodology and inference speedups of Granite 5.0 Turbo CTC, a 470 million parameter encoder-only model with an excellent speed-accuracy tradeoff. The architecture uses pyramidal temporal subsampling within Conformer blocks using strided depthwise convolutions, block-diagonal (chunk-wise) self-attention, and conditioning on intermediate predictions from the middle layer. Training highlights are the use of only publicly available data, the novel use of a Muon optimizer, and balanced data sampling. Inference speedups include replacing 1 x 1 convolutions with linear layers and optimizing the attention computation in the Conformer blocks. Collectively, these result in a model that is on the speed-accuracy Pareto frontier of the Open ASR leaderboard for English short-form ASR while being twice as fast as the fastest competitor. The model can be used under a permissive license and downloaded from https://huggingface.co/ibm-granite/granite-speech-5.0-470m-turboctc.

Comments5 pages, 2 figures, submitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑