Orukeet:具有冻结Gabor核的多语言自动语音识别
Orukeet: Multilingual ASR with Frozen Gabor Kernels
- Oruk AI
- Stanford University(斯坦福大学)
- University of Cambridge(剑桥大学)
- OpenWhispr
- Hoid
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Orukeet用冻结的Gabor核替换Parakeet一半时间滤波器,在多语言数据上训练,使25种语言平均WER从11.01%降至9.85%,相对降低10.6%,并在多数子集上表现更优。
AI中文摘要:
Orukeet将经过适配的Parakeet编码器中一半的时间滤波器替换为12,288个拟合的Gabor核,冻结这些替换部分,并在多语言和多口音数据上训练其余参数。最终适配和检查点选择使用LibriSpeech test-other。在25种语言的20,146条FLEURS录音中,合并词错误率(WER)从Parakeet的11.01%降至Orukeet的9.85%,相对降低10.6%。Orukeet在25种语言中的23种上取得了更低的WER。在74个测试子集中,Orukeet在61个上优于Parakeet,包括LibriSpeech test-clean(1.46%对1.53%的WER)、test-other(2.86%对3.14%)以及FLEURS英语(3.82%对4.28%)。所有比较均在匹配的NeMo设置下解码相同的音频。拟合的核以普通卷积权重存储,保留了Parakeet的架构和推理算子。
英文摘要:
Orukeet replaces half of an adapted Parakeet encoder's temporal filters with 12,288 fitted Gabor kernels, freezes these replacements, and trains the remaining parameters on multilingual and multi-accent data. Final adaptation and checkpoint selection use LibriSpeech test-other. Across 20,146 FLEURS recordings in 25 languages, pooled word error rate (WER) falls from Parakeet's 11.01% to Orukeet's 9.85%, a 10.6% relative reduction. Orukeet has lower WER on 23 of the 25 languages. Orukeet outperforms Parakeet on 61 out of 74 tested splits, including LibriSpeech test-clean (1.46% vs. 1.53% WER), test-other (2.86% vs. 3.14%), and FLEURS English (3.82% vs. 4.28%). All comparisons decode the same audio with matched NeMo settings. The fitted kernels are stored as ordinary convolution weights, retaining Parakeet's architecture and inference operators.