arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10054cs.SDcs.LGeess.AS

Orukeet:具有冻结Gabor核的多语言自动语音识别

Orukeet: Multilingual ASR with Frozen Gabor Kernels

  • Oruk AI
  • Stanford University(斯坦福大学)
  • University of Cambridge(剑桥大学)
  • OpenWhispr
  • Hoid

机构由 AI 辅助整理,请以论文原文为准。

Nathan Roll, Irene Yi, Büşra Marşan, Vianney Grenez, Gabriel Stein, Momcilo Mrkaic, Pavle Padjin, Vladimir Zeljkovic, Calbert Graham

AI总结:

Orukeet用冻结的Gabor核替换Parakeet一半时间滤波器,在多语言数据上训练,使25种语言平均WER从11.01%降至9.85%,相对降低10.6%,并在多数子集上表现更优。

AI中文摘要:

Orukeet将经过适配的Parakeet编码器中一半的时间滤波器替换为12,288个拟合的Gabor核,冻结这些替换部分,并在多语言和多口音数据上训练其余参数。最终适配和检查点选择使用LibriSpeech test-other。在25种语言的20,146条FLEURS录音中,合并词错误率(WER)从Parakeet的11.01%降至Orukeet的9.85%,相对降低10.6%。Orukeet在25种语言中的23种上取得了更低的WER。在74个测试子集中,Orukeet在61个上优于Parakeet,包括LibriSpeech test-clean(1.46%对1.53%的WER)、test-other(2.86%对3.14%)以及FLEURS英语(3.82%对4.28%)。所有比较均在匹配的NeMo设置下解码相同的音频。拟合的核以普通卷积权重存储,保留了Parakeet的架构和推理算子。

英文摘要:

Orukeet replaces half of an adapted Parakeet encoder's temporal filters with 12,288 fitted Gabor kernels, freezes these replacements, and trains the remaining parameters on multilingual and multi-accent data. Final adaptation and checkpoint selection use LibriSpeech test-other. Across 20,146 FLEURS recordings in 25 languages, pooled word error rate (WER) falls from Parakeet's 11.01% to Orukeet's 9.85%, a 10.6% relative reduction. Orukeet has lower WER on 23 of the 25 languages. Orukeet outperforms Parakeet on 61 out of 74 tested splits, including LibriSpeech test-clean (1.46% vs. 1.53% WER), test-other (2.86% vs. 3.14%), and FLEURS English (3.82% vs. 4.28%). All comparisons decode the same audio with matched NeMo settings. The fitted kernels are stored as ordinary convolution weights, retaining Parakeet's architecture and inference operators.

补充信息

↑