arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SraVaani 1.0:面向印度语言的包容性自动语音识别规模化模型

SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages

Sujith Pulikodan, Agneedh Basu, Pavan Kumar J, Pranav D Bhat, Suryansh Shukla, Nihar Desai, Prasanta Kumar Ghosh

arXiv 2608.08235首次发表:更新:

AI 中文总结

本文提出SraVaani 1.0,一款基于FastConformer架构的多语言ASR模型,覆盖65种印度语言,经三阶段训练后在8个基准上表现优异,是唯一支持多种低资源及部落印度语言转录的开源模型。

AI 中文摘要

印度的语言格局涵盖700余种语言与数千种方言,但绝大多数自动语音识别(ASR)系统仅支持其中极小部分语言。本文提出SraVaani-1.0,一款覆盖65种印度语言及方言的多语言ASR模型,其中多数语言目前无公开可用或有竞争力的ASR系统。SraVaani-1.0基于FastConformer架构构建,通过三阶段流程从头训练:第一阶段,采用对比学习目标,对VAANI语料库中31255小时未标注语音进行自监督预训练;第二阶段,引入音频-图像表示对齐阶段,利用VAANI语料库中配对的图像与语音,该多模态对齐通过挖掘视觉上下文与语音内容的关系,促使语音编码器学习语义更丰富的表示,从而提升下游识别性能,尤其针对低资源语言;最后阶段,使用混合令牌-时长换能器(TDT)-CTC解码器,对从24个公开数据集汇编的、覆盖65种语言及方言的31263小时标注多语言印度语音,对已对齐的编码器进行端到端微调。我们在8个基准上将SraVaani-1.0与3款最先进的多语言ASR系统对比评估,结果显示SraVaani-1.0在大量语言-数据集对中取得最低词错误率(WER),同时在高资源语言上与性能最优的系统具有竞争力;重要的是,它是唯一经评估的开源模型,能为多种低资源及部落印度语言提供转录能力,这些语言仅在VAANI基准上进行评估。

英文摘要

India's linguistic landscape spans over 700 languages and thousands of dialects, yet the vast majority of automatic speech recognition (ASR) systems support only a small fraction of this diversity. We present SraVaani-1.0, a multilingual ASR model covering 65 Indian languages and dialects, many of which currently have no publicly available or competing ASR system. SraVaani-1.0 is built on a FastConformer architecture and trained from scratch through a three stage the first stage, we perform self-supervised pretraining on 31,255 hours of unlabelled speech from the VAANI corpus using a contrastive learning objective. In the second stage, we introduce an audio-image representation alignment stage that leverages the paired images and speech available in the VAANI corpus. This multimodal alignment encourages the speech encoder to learn semantically richer representations by exploiting the relationship between visual context and spoken content, thereby improving downstream recognition, particularly for low resource the final stage, the aligned encoder is fine-tuned end-to-end using a Hybrid Token-and-Duration Transducer (TDT)-CTC decoder on 31,263 hours of labelled multilingual Indian speech compiled from 24 public datasets spanning 65 languages and dialects. We evaluate SraVaani-1.0 against three state-of-the-art multilingual ASR systems across eight benchmarks. SraVaani-1.0 achieves the lowest word error rate (WER) on a large number of language-dataset pairs while remaining competitive with the best-performing systems on high resource importantly, it is the only open-source evaluated model that provides transcription capability for multiple low-resource and tribal Indian languages, which are assessed exclusively on the VAANI benchmark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑