arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Vimarsha:面向印度语言的高保真ASR评估,涵盖人口多样性、野外音频与拼写变体

Vimarsha: Faithful ASR Evaluation for Indian Languages with Demographic Diversity, In-the-Wild Audio and Spelling Variations

Kaushal Santosh Bhogale, Srija Anand, Sadakopa Ramakrishnan Thothathiri, Tahir Javed, Sshubam Verma, Mitesh M. Khapra

arXiv 2609.24199首次发表:更新:

发表机构

Indian Institute of Technology, Madras; Sarvam AI(印度理工学院马德拉斯分校; Sarvam AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对印度语言ASR评估的偏差,提出100小时覆盖22种语言的Vimarsha基准,结合人口多样性与野外音频及变体格框架,评估10个模型发现排名显著变化和系统性失败模式。

AI 中文摘要

印度语言自动语音识别(ASR)的评估基准存在两个系统性偏差:一是来自干净、受控音频条件的乐观分数,二是来自过于僵化的转录标准所导致的悲观分数,这些标准会惩罚有效的语言变体。我们推出了Vimarsha,一个覆盖全部22种印度法定语言的100小时基准,旨在解决这两种失真。Vimarsha结合了人口多样性的实地录音和经过精心挖掘、按声学难度筛选的野外音频,并采用变体格框架,为每个话语编码多个有效的转录。对10个最先进的ASR模型的评估显示,在现实条件下模型排名发生显著变化,存在地理和人口统计上的性能差异,以及在不同语速和声学环境中的系统性失败模式。

英文摘要

Evaluation benchmarks for Indian language automatic speech recognition (ASR) suffer from two systematic biases: optimistic scores from clean, controlled audio conditions, and pessimistic scores from overly rigid transcription standards that penalize valid linguistic variations. We introduce Vimarsha, a 100-hour benchmark spanning all 22 scheduled Indian languages, designed to address both distortions. Vimarsha combines demographically diverse on-field recordings with carefully mined in-the-wild audio selected for acoustic difficulty, alongside a lattice of variations framework that encodes multiple valid transcriptions per utterance. Evaluations of 10 state-of-the-art ASR models reveal substantial shifts in model rankings under realistic conditions, geographic and demographic performance disparities, and systematic failure modes across speaking rates and acoustic environments.

CommentsAccepted in Interspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑